Claude Opus 5.5: Fable-level work at a lower Opus price
Anthropic's new mid-priced model scores at or above its most expensive one on most published tests, and it costs less than the Opus it replaces. Here is what that means in practice, where it falls short and how to try it on work you already do.
By Monolith

Picture a small agency looking after a client's booking system. It was built six years ago, nobody who wrote it still works there, and the owner wants online payments added before the busy season. The job is part archaeology and part surgery: read a large, unfamiliar codebase, change it without breaking the parts customers rely on, then prove the change works. That is exactly the kind of work Claude Opus 5.5 is aimed at.
Anthropic released Opus 5.5 on September 22, 2026, two months after Opus 5. The short version: it performs at roughly the level of Claude Fable 5.1, Anthropic's larger and pricier model, on most of the company's published tests, and it costs less per token than the Opus 5 it replaces. Anthropic says that at default settings a typical workload costs about 40% less than on Opus 5, and that output arrives about 30% faster. 1
Where Opus 5.5 sits in the Claude lineup
Anthropic sells several model sizes at once. Bigger models tend to handle harder problems but cost more and answer more slowly. Anthropic's own documentation now tells developers to start with Opus 5.5 for most workloads, and to move up to Fable 5.1 only when Opus 5.5 at a higher effort setting still falls short on their tests. 2
| Model | Built for | Price | Speed | Context window |
|---|---|---|---|---|
| Claude Fable 5.1 | The hardest reasoning and very long agent tasks | $10 / $50 | Slower | 1M tokens |
| Claude Opus 5.5 | Long-running coding and knowledge work | $4 / $20 | Moderate | 1M tokens |
| Claude Sonnet 5 | A balance of speed and intelligence | $2 / $10 | Fast | 1M tokens |
| Claude Haiku 4.5 | Quick, high-volume tasks | $1 / $5 | Fastest | 200K tokens |
A token is a small piece of text; on Claude's current tokenizer it averages a little over half an English word. Prices are quoted per million of them. The context window is how much material the model can consider in one request: 1 million tokens is roughly 555,000 words, enough for a large codebase or a year of project documents. Opus 5.5 can write up to 128,000 tokens in a single reply, and its reliable knowledge runs to June 2026. Anthropic says smaller Sonnet 5.5 and Haiku 5.5 models are coming in the following weeks. 12
What the coding benchmarks say
A benchmark is a fixed set of tasks used to compare systems under the same conditions. The three below measure agentic coding: the model works through a real software task using tools such as a terminal, rather than answering a single question. Terminal-Bench 4.0 sets jobs inside a command line, FrontierCode tests difficult programming problems end to end, and CursorBench comes from Cursor, a popular AI code editor. All figures are Anthropic's own published results. 1
The jump on Terminal-Bench is the one to notice: about 14 points over Opus 5 and 10 over Fable 5.1. On FrontierCode the lead over OpenAI's GPT-6 Astra is only about one point, which is within the range where a different test setup could flip the order. Anthropic also reports that at default effort Opus 5.5 beats GPT-6 Astra on FrontierCode at roughly a fifth of the cost per task. That is a vendor comparison, and the cost side depends heavily on how each model is configured. 1
Anthropic's examples point the same way. One tester finished a 680,000-line code migration in less than a day. Asked to cut load times across every page of a web app, Opus 5.5 succeeded 39 times out of 40, while Opus 5 made smaller improvements that also changed how the app behaved. Those are selected examples, not averages, but they describe the booking-system job well: large, unfamiliar code where a fix must not quietly break something else. 1

Beyond code: research, office work and computer use
Most businesses will not use Opus 5.5 to write software. The broader results matter more for them. Humanity's Last Exam is a set of very hard expert questions across many subjects. OSWorld 2.0 measures whether a model can operate a computer through its screen, clicking and typing like a person. Chartography tests reading values off charts. AutomationBench looks at multi-step business workflows, and Terminal-Bench-Science at scientific research tasks. 1
Read the last two rows carefully. GPT-6 Astra is ahead on business workflow automation, narrowly, and clearly ahead on the scientific research test. A company choosing a model for automated research pipelines should not take the headline claim at face value; it should compare both models on its own tasks. Also notice how close Opus 5.5 and Fable 5.1 are on computer use and chart reading. For those jobs, the cheaper model looks like the sensible default.
Knowledge work is scored differently. GDPval-AA pits models against each other on realistic professional tasks, such as drafting a memo or building a spreadsheet, and ranks them with an Elo rating, the same kind of score used for chess players. The number only means something relative to the other models in the same table. 1
| Model | Elo rating | Gap to Opus 5.5 |
|---|---|---|
| Claude Opus 5.5 | 1846 | – |
| Claude Fable 5.1 | 1735 | 111 lower |
| Claude Opus 5 | 1708 | 138 lower |
| GPT-5.6 Sol | 1588 | 258 lower |
| GPT-6 Astra | 1542 | 304 lower |
The price cut, and why the real saving can be bigger
Every line of the price list came down. Input fell from $5 to $4 per million tokens and output from $25 to $20. Reading from the prompt cache, which is reused material such as a long set of instructions or a codebase the model has already seen, dropped 60%, from $0.50 to $0.20. Batch requests, which run in the background instead of immediately, stay at half price. 13
| Price line | Opus 5 | Opus 5.5 | Fable 5.1 |
|---|---|---|---|
| Input | $5 | $4 | $10 |
| Output | $25 | $20 | $50 |
| Cache read | $0.50 | $0.20 | $0.25 |
| 5-minute cache write | $6.25 | $5 | – |
| Batch input / output | $2.50 / $12.50 | $2 / $10 | $5 / $25 |
The list price is only half of it. Anthropic says Opus 5.5 writes more directly and puts the important information first, and several launch customers report using fewer tokens for the same work: Factory cites 20 to 25% fewer output tokens and Kiro 40% fewer calls than Opus 5. Fewer words at a lower price per word is where the 40% figure comes from. 1
Here is a worked example. Imagine one long session on the booking system that reads 2 million tokens, 1.5 million of them from the cache because the same code is read again and again, plus 500,000 fresh tokens, and writes 200,000 tokens of changes and notes. We have left cache writes out to keep the arithmetic simple.
Opus 5.5 comes out about 24% cheaper than Opus 5 before counting any reduction in tokens, and well under half the cost of Fable 5.1. If it also finishes with a fifth fewer output tokens, as Factory reports, the gap widens further. Your own numbers depend on how much you reuse, how long the replies are and which effort level you choose.

Fast mode: paying for speed
Fast mode runs the same model on a faster serving configuration. Anthropic says it produces output up to 2.5 times faster, with no change to the model's intelligence. It costs double: $8 per million input tokens and $40 per million output. It is available in Claude Code and on the Claude API, where it is a research preview that requires access, and it is not offered on Amazon Bedrock, Google Cloud or Microsoft Foundry. 14
The speed-up applies to how quickly the answer is written, not how quickly it starts. That makes fast mode worth it where someone is sitting and waiting, such as a live pairing session or a demo with a client in the room. For overnight or background jobs, standard speed or the half-price batch option usually makes more sense. 4
Effort: the dial that now matters most
Opus 5.5 always thinks before it answers; that can no longer be switched off. What you control is effort, which runs from low to max and sets how much thinking the model does. The default is now medium, where Opus 5 defaulted to high. Anthropic also notes that at the same setting Opus 5.5 tends to think more per turn than Opus 5, especially at the top levels. 3
In practice: medium effort for routine drafting and edits, high for complex changes, and the top settings only for problems that have already defeated a normal attempt. More effort means more tokens, so the dial is also a cost control. If a team copies its old Opus 5 settings across without re-testing, it may pay for thinking it does not need, or get shallower work than before.
What developers need to change
For people using Claude through an app, nothing needs to change. Teams that call the model from their own software should read Anthropic's migration notes, because four changes can make existing requests fail with an error, and a fifth changes what a response looks like. 3
| Change | What happens | What to do |
|---|---|---|
| Thinking is always on | Requests that disable thinking or set a manual budget return an error | Remove those settings and choose an effort level |
| No forced tool use | Requiring a specific tool returns an error | Use automatic tool choice with strict tool use, or structured outputs |
| Older computer use tool retired | The earlier computer use tool is rejected on the Claude API and Google Cloud | Move to the newer computer use toolset |
| Thinking tied to the conversation | Reasoning from Fable or Mythos models is dropped when a chat moves to Opus 5.5 | Keep conversations append-only and test model switching |
| Progress notes moved | Short notes between tool calls arrive as thinking blocks, hidden by default | Change the display setting if your app shows progress updates |
Safeguards, and the requests that go to another model
Anthropic says Opus 5.5 has very strong cybersecurity abilities, so it carries safeguards similar to Fable 5.1's. Routine work, such as finding and fixing bugs in your own code, is allowed. Most other cybersecurity tasks are re-routed to the older Claude Opus 4.8, and a verification program is expanding for security professionals who need more. The model also runs a biology safety classifier. 13
For a business, the practical point is that a small number of requests may be declined or answered by a different model than the one you chose. On the API, a declined request is labelled as a refusal and names the policy area, and there is an optional fallback that retries on the model Anthropic recommends. A security consultancy, or anyone running automated security scanning, should test its own tasks before relying on Opus 5.5 for them. 3

Anthropic also reports that Opus 5.5 was its best-performing model on its automated behavior audit, that it is 85% less likely than Opus 5 to try to get around containment boundaries in testing, and that it resists prompt injection, meaning instructions hidden in web pages or documents, across coding, browsing and computer use. These are the developer's own safety findings. They reduce risk; they do not remove the need to limit what an agent can access. 1
What early customers reported
Launch announcements always include customer quotes. They are useful as a sign of where a model helped, but they come from companies that had early access and were chosen by the vendor. We list a few that include a measurable claim.
| Company | What they reported | Kind of work |
|---|---|---|
| Deloitte | Caught 72% of bugs in their test, versus 56% for Opus 5 | Code review |
| Hebbia | 86.6% coverage on finance workflows, versus 60.3% for Opus 5 | Financial research |
| Optiver | 40 to 50% lower cost on agentic coding | Software development |
| Kiro | 40% fewer calls than Opus 5 | Coding agent |
| Factory | 20 to 25% fewer output tokens | Coding agent |
Should you switch?
For most teams already on Opus 5, yes, after a short test. The model is cheaper and scores better on nearly every published measure. The questions are narrower: which tasks still need Fable 5.1, which should drop down to Sonnet, and whether any of your work touches the security categories that get re-routed.
| If your work is mostly... | Start with | Why |
|---|---|---|
| Large code changes, audits or migrations | Opus 5.5 | Leads the published coding tests at a mid-range price |
| Reports, research and document-heavy work | Opus 5.5 | Highest published knowledge-work rating |
| Problems Opus 5.5 at high effort still fails | Fable 5.1 | Anthropic's recommended step up |
| High-volume drafting, sorting and replies | Sonnet 5 or Haiku 4.5 | Much cheaper and faster when the task is simple |
| Automated scientific research pipelines | Test Opus 5.5 against GPT-6 Astra | Astra leads the published science test |
| Security testing beyond fixing your own bugs | Check eligibility first | Most such tasks route to Opus 4.8 |
How to run a fair one-week trial
Go back to the booking system. Rather than switching everything on day one, pick three tasks your team did recently and still remembers well: one routine, one tricky and one that went badly the first time. Run each on the old setup and on Opus 5.5 with the same brief, and compare what you actually care about.
- Choose three real, recent tasks with a known good outcome.
- Write one brief per task and use it unchanged for both models.
- Set effort explicitly rather than relying on the new default.
- Record corrections needed, review time and total cost for each run.
- Check the work the way a client would: open the page, run the tests, read the report.
- Keep the model that produces accepted work at the lowest total cost.

Keep a person in the loop wherever a change reaches customers or money. Launch customers describe long unattended runs with Opus 5.5, one of them lasting more than 18 hours. That makes the checks at the end more important, not less. A long run that finishes with a clear list of what was changed, what was tested and what could not be verified is worth far more than one that simply says it is done. 1
Monolith's take: Opus 5.5 is the new default for serious work on Claude. It costs less than Opus 5, matches or beats Fable 5.1 on most published tests, and writes its answers faster. Trial it on three real tasks, set effort on purpose, and keep Fable and the smaller models for the jobs where your own results say they fit better.
Is Claude Opus 5.5 better than Claude Fable 5.1?
How much does Claude Opus 5.5 cost?
Read for this feature. The numbers match the markers in the text.
- Anthropic: Introducing Claude Opus 5.5, benchmarks and customer statements, September 22, 2026anthropic.com
- Anthropic docs: Models overview; checked September 22, 2026platform.claude.com
- Anthropic docs: What's new in Claude Opus 5.5, breaking changes and pricingplatform.claude.com
- Anthropic docs: Fast mode (research preview) pricing and availabilityplatform.claude.com
Want your team to run this kind of model trial with confidence? Bring a real task to a workshop and leave with a repeatable process.
Want to apply this to a project? Here is the related service.
25 AI instructions to try on everyday work
Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.