InsightsPublished 13 min read

Claude Haiku 5.5: Anthropic's cheapest model just got a lot smarter

Anthropic's small, fast model now costs a tenth of what Haiku 4.5 did for most requests, and on Anthropic's own tests it closes much of the gap to Sonnet. Here is what it costs, what the independent numbers say, and the high-volume jobs it is built for.

By Monolith

A support specialist typing on a laptop at a standing desk in an outdoor gear shop while a colleague packs online orders behind her
A support specialist answering customer messages while a colleague packs the day's online orders.

Picture an outdoor gear shop that sells as much online as it does over the counter. Every day brings a few hundred customer messages: where is my order, does this tent fit two, can I swap the size. Most need sorting, a short summary and a reply drafted from the shop's own policies, and only a few need a person. That kind of steady, repetitive, high-volume work is what small AI models are for, and it is where cost per message decides whether automation pays for itself.

Anthropic released Claude Haiku 5.5 on October 7, 2026, joining Opus 5.5 and Sonnet 5.5 in its Claude 5.5 family. Anthropic calls it the cheapest, fastest and most capable small model it has released, built for high-volume, cost-sensitive tasks such as summaries, database queries and classification, for speed-sensitive work such as live customer support and browser use, and as a helper to the larger models on coding jobs. Anthropic says it costs around 75% less to run than Haiku 4.5. 1

What Haiku 5.5 is built for

Anthropic sells its models in sizes. Opus and Sonnet handle long, complicated work; Haiku is the small one, meant for jobs where you run the same kind of task thousands of times and need each answer quickly and cheaply. Anthropic's documentation describes Haiku 5.5 as built for classification, routing, extraction and subagent tasks, and lists it as the fastest model in the current line-up. 2

The subagent role is worth understanding. In an agent setup, a larger model plans the work and hands narrow pieces to smaller, cheaper helpers: look up one figure in a long report, summarize a thread, check a record. Anthropic says Haiku 5.5 pairs well with Opus 5.5 and Sonnet 5.5 in exactly that way. It is also, by Anthropic's account, its fastest model at standard speed, though Opus models running in fast mode can still be quicker. 1

ModelPriceSpeedDefault effortContext window
Claude Fable 5.1$10 / $50SlowerHigh1M tokens
Claude Opus 5.5$4 / $20ModerateMedium1M tokens
Claude Sonnet 5.5$2 / $10FastHigh1M tokens
Claude Haiku 5.5From $0.10 / $0.50FastestMedium1M tokens
Current Claude models on the API, from Anthropic's Claude Haiku 5.5 documentation; checked October 7, 2026. Prices are USD per million tokens, input / output; Haiku 5.5's lower price applies to prompts up to 100,000 tokens. 2

The price: a tenth of Haiku 4.5 for most requests

Haiku 5.5 has two price rows, set by the length of the prompt. For prompts up to 100,000 tokens, input costs $0.10 per million tokens and output $0.50, a tenth of Haiku 4.5's $1 and $5. Above 100,000 tokens the price rises to $0.50 and $2.50, which is still half of Haiku 4.5. Anthropic says about 90% of requests to Haiku 4.5 fell in the shorter band. Batch requests, which run in the background, get a further 50% off. 12

Price lineHaiku 5.5Haiku 4.5Sonnet 5.5
Input$0.10 / $0.50$1.00$2.00
Output$0.50 / $2.50$5.00$10.00
Cache read$0.01 / $0.05$0.10$0.10
5-minute cache write$0.125 / $0.625$1.25$2.50
Claude API list prices in USD per million tokens, from Anthropic's Haiku 5.5 announcement and documentation; checked October 7, 2026. Haiku 5.5 shows prompts up to / over 100,000 tokens. Sonnet 5.5's cache read price reflects the cut announced the same day. 12

Why 75% and not 90%? Haiku 5.5 uses Anthropic's newer tokenizer, the same one as Sonnet 5.5 and Opus 5.5, so the same text counts as roughly 30% more tokens than it did on Haiku 4.5. A token is a small piece of text, and you pay per token, so some of the price cut is given back in token count. Anthropic's 75% figure accounts for that, and for the share of requests over 100,000 tokens. 12

Here is a worked example for the gear shop. Say it runs 100,000 support messages a month through a model that reads each one with its policies and writes a short draft: about 3,000 tokens in and 300 out, as counted on Haiku 4.5. On the newer tokenizer that becomes about 3,900 in and 390 out.

One caution on that figure. Haiku 5.5 thinks before it answers when the task calls for it, and thinking tokens count toward its output. Even if thinking tripled the output tokens, the month would come to about $98, still under a quarter of the Haiku 4.5 figure. The lever for that is the effort setting, covered below. 24

What Anthropic's benchmarks show

Anthropic published results against Haiku 4.5, OpenAI's budget GPT-6 Luna and its own Sonnet 5.5 for reference. The jump from Haiku 4.5 is large on every row. OSWorld 2.1, which tests whether an agent can operate a real computer through long multi-step tasks, goes from 15.7% to 72.4%. Terminal-Bench 4.0, a test of complex jobs in a command line, goes from zero to 39.2%. 1

Two readings matter. Against GPT-6 Luna, which costs the same per token, Haiku 5.5 leads on every test Anthropic listed for both, by more than 20 points on computer use. Against Sonnet 5.5, it trails, and by far the most on Terminal-Bench 4.0, where Sonnet scores 70.6%. Anthropic is direct about this: Sonnet and Opus remain the better choice for complex agentic coding, and Haiku is best for narrower jobs such as summarizing, compacting long conversations or subagent work. 1

Knowledge work is scored with Elo ratings, the same kind of score used to rank chess players, from two Artificial Analysis tests of professional tasks such as drafting documents and building spreadsheets. The ratings only mean something relative to the other models in the same table. 1

TestHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1162073514371840
AA-Briefcase v1.1157861413361824
Elo ratings as published in Anthropic's announcement, checked October 7, 2026. Higher is better; scores are comparable only within each row. 1

These are Anthropic's own figures, run its way and chosen for its launch. The independent results below run every model through the same tests.

What independent testing says

Artificial Analysis scores Haiku 5.5 at maximum effort at 43 on its Intelligence Index, which combines ten evaluations. That ranks it second of the 182 models Artificial Analysis groups with it by price, where the median is 13. GPT-6 Luna at maximum effort scores 38. The larger models still sit well above: GPT-6.1 Sol at 52 and Claude Opus 5.5 at 58. 3

The chart shows the trade-off clearly. At the same score, GPT-6 Luna is slightly cheaper per task: Luna at maximum effort scores 38 for about $0.07, where Haiku 5.5 at high effort scores 38 for about $0.08. But Haiku keeps climbing past where Luna stops, to 43 at about $0.21 per task, still less than a third of GPT-6.1 Sol's $0.72. 3

Speed is where Haiku stands out. Artificial Analysis measured it at 243 output tokens per second at maximum effort, ninth fastest of the 182 models in its group, roughly twice GPT-6 Luna and more than four times GPT-6.1 Sol. For live chat, where a customer is watching the reply appear, that difference is visible. 3

The third result matters most for anything customer-facing. On Artificial Analysis's AA-Omniscience test, which asks hard knowledge questions and penalizes confident wrong answers, Haiku 5.5's hallucination rate is 40%, against 77% for GPT-6 Luna. When it does not know, it is far more likely to say so. It also knows less than the bigger models: it answers 36% of the questions correctly. For support work, that is the right trade, as long as the model answers from your own documents rather than from memory. 3

Effort: the first Haiku with a dial

Haiku 5.5 is the first Haiku model with an adjustable effort setting, running from low to max. Effort controls how much the model thinks before answering, which trades cost and speed against quality. The default is medium. Thinking can no longer be given a fixed token budget; instead the model decides how much to think, and effort steers it. 124

The setting changes the bill more than the price list suggests. In Artificial Analysis's runs, Haiku 5.5 at medium effort used about 33,000 output tokens per index task and scored 34; at maximum effort it used about 162,000 and scored 43. Artificial Analysis calls the model very verbose at maximum effort. For sorting and routing messages, start at low or medium and raise effort only for the steps where quality falls short. 3

What early customers reported

Launch customer statements come from companies with early access, chosen by the vendor. They are useful as a sign of where the model helped. These are the ones with a measurable claim.

CompanyWhat they reportedKind of work
AsanaOver 30% lower latency on task completions, up to 2.5x faster per agent turnProject management agents
HubSpot92.8% on its CRM test suite, the best score it has seen thereCRM reporting and audits
AlphaSense0.84 versus 0.76 for Haiku 4.5 across 400 queriesQuestions over documents
Box11 points higher than Haiku 4.5 at about half the latencyAnalysis of enterprise content
RogoTrusted as a subagent pulling figures from financial filingsFinancial research
Customer statements as published in Anthropic's announcement, October 7, 2026. Not independently verified by Monolith. 1
An operations manager and a dispatcher on the phone looking at a monitor together in a home-services office with service vans outside
An operations manager and a dispatcher working through the morning's service requests.

Where it fits in a small business

For most small businesses, Haiku 5.5 is not the model you chat with. It is the one working behind a form, an inbox or a help widget, doing the same small job over and over. For support that answers from your own documents, it is now a strong, low-cost default for the answering step.

If the job is...Start withWhy
Sorting, tagging or routing incoming messagesHaiku 5.5, low effortFast and very cheap per message
Live chat answered from your own policiesHaiku 5.5, medium effortFastest Claude model with a low hallucination rate
Summaries of calls, threads and long documentsHaiku 5.5Anthropic's named use case, 1M token context
Filling forms or checking sites in a browserHaiku 5.5, then testBig jump on computer use; review before anything submits
Multi-step coding or complex agent workSonnet 5.5 or Opus 5.5Haiku trails badly on Terminal-Bench 4.0
Very large prompts, over 100,000 tokensCompare Haiku and SonnetHaiku's price rises fivefold above that line
Monolith's suggested starting points, based on Anthropic's published results and Artificial Analysis's measurements. Suggestions, not results of Monolith testing. 13

Keep a person on anything that sends money, changes an order or makes a promise to a customer. A fast, cheap model makes it tempting to automate everything; the sensible pattern is to let it draft and sort, and have someone approve the steps that cannot be undone.

Building support that answers from your own help pages and policies? Our guide walks through the setup, the guardrails and what it costs.

Also announced: cheaper Sonnet and API credits for subscribers

Anthropic cut the price of cache reads on Sonnet 5.5 in half, from $0.20 to $0.10 per million tokens. Cache reads are the cost of re-reading material the model has already seen, such as long instructions or a codebase, and they make up a large share of agent work, so Anthropic says Sonnet 5.5 now runs about 20% cheaper on most agentic tasks. 1

Subscribers also get API money to build with. This week Anthropic is rolling out a monthly API credit for the Claude Platform: $100 for Max 5x subscribers, $200 for Max 20x, and up to $500 for Team plans, pooled across users. The credits work on any model. For a small team already paying for Claude, that is enough to prototype a support or sorting workflow on Haiku 5.5 without a separate budget. Anthropic is also adding computer use and browser use to its Python and TypeScript SDKs in beta. 1

What developers need to change

Nothing changes for people using Claude through an app. Teams calling Haiku 4.5 from their own software should read Anthropic's migration guide first, because several old request settings now fail with an error. 4

ChangeWhat happensWhat to do
New model IDUse claude-haiku-5-5; it is fixed, with no date suffix or separate aliasUpdate the ID on each platform you use
About 30% more tokensUsage counts rise and old max_tokens limits may cut replies shortRecount prompts and redo cost estimates
Thinking budgets removedA fixed thinking budget returns an errorUse adaptive thinking and choose an effort level
Sampling settings removedCustom temperature, top_p or top_k returns an errorRemove them and steer with the prompt
No assistant prefillEnding a request with a partial assistant reply returns an errorUse structured outputs or move it to the user turn
New computer use toolsetThe older computer use tool is rejected on the Claude API and Google CloudMove to the newer toolset
Refusals have no fallbackSafety classifiers can decline a request, with no automatic retry on another modelHandle refusals in your code
Changes when moving from Claude Haiku 4.5, summarized from Anthropic's migration guide; checked October 7, 2026. 4

Two more points for planning. Priority Tier, Anthropic's committed-capacity service level, is not supported on Haiku 5.5, so teams relying on it for Haiku 4.5 should plan capacity separately. And the model's cybersecurity safeguards are stricter than Haiku 4.5's: defensive work is allowed, but penetration testing and other techniques more likely to be used by attackers are blocked. Haiku 4.5 remains available for now. 14

Two developers talking at a tall table with laptops in a brick-walled studio while a colleague works with headphones behind them
Two developers planning a model switch in a small software studio.

How to test it in a week

Back to the gear shop. Rather than switching everything at once, take a sample of last month's real messages, including the awkward ones, and run them through the current setup and through Haiku 5.5 with the same instructions. Then compare what matters to the business, not the benchmark.

  • Pull 200 recent messages with known good outcomes, including edge cases.
  • Run them on your current model and on Haiku 5.5 with identical instructions.
  • Try low and medium effort, and record cost and response time for each.
  • Count wrong answers and wrong routings, not just how good the replies sound.
  • Check that answers come from your documents, and that it says when it does not know.
  • Switch only the tasks where Haiku 5.5 is as accurate at a lower total cost.

Comparing the bigger Claude models for heavier work? Our Opus 5.5 guide covers the prices, benchmarks and effort settings.

The short version

Monolith's take: Claude Haiku 5.5 is the new default for high-volume, low-stakes AI work. It costs a fraction of Haiku 4.5, it is the fastest Claude model, and on independent testing it is both smarter and far less likely to make things up than GPT-6 Luna at the same token price. Use it for sorting, summaries and support drafts, keep Sonnet or Opus for complex agent and coding work, and set effort on purpose.

Questions, answered
How much does Claude Haiku 5.5 cost?
For prompts up to 100,000 tokens, $0.10 per million input tokens and $0.50 per million output tokens. Longer prompts cost $0.50 and $2.50. Batch requests are half price.
Is Claude Haiku 5.5 better than GPT-6 Luna?
On Anthropic's published tests and on Artificial Analysis's Intelligence Index it scores higher, and its hallucination rate is far lower. At the same score, Luna is slightly cheaper per task.
Should I use Haiku 5.5 or Sonnet 5.5?
Use Haiku 5.5 for narrow, repeated tasks such as sorting, summaries and support replies. Use Sonnet 5.5 or Opus 5.5 for complex coding and multi-step agent work, where Haiku trails.
Sources

Read for this feature. The numbers match the markers in the text.

  1. Anthropic: Introducing Claude Haiku 5.5, benchmarks, pricing and customer statements, October 7, 2026anthropic.com
  2. Anthropic docs: Claude Haiku 5.5 overview, specifications and pricingplatform.claude.com
  3. Artificial Analysis: Claude Haiku 5.5artificialanalysis.ai
  4. Anthropic docs: Claude Haiku 5.5 migration guideplatform.claude.com

Want AI sorting your inbox, drafting support replies or handling the repetitive steps in your workflow? That is what we build.

Where this leads

Want to apply this to a project? Here is the related service.

Free download

25 AI instructions to try on everyday work

Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.

Enter your email to access the PDF. This form does not sign you up for a marketing sequence.