Guides
Claude Opus 5.5 vs Sonnet 5.5: Which Should You Use? (The Cheaper One Isn't Always Cheaper)
By the Top5Apps editorial team · Published September 29, 2026 · Updated September 29, 2026 · 6 min read
Short answer: use Sonnet 5.5 at medium or high effort for well-defined work, and switch to Opus 5.5 the moment you're tempted to turn Sonnet up to max. Sonnet 5.5 launched September 28 at $2 in / $10 out per million tokens — half of Opus 5.5's $4/$20 — and it's faster. But list price isn't what you pay. At max effort, Sonnet 5.5 uses so many tokens that independent testing found it costs more per task than Opus ($7.60 vs $5.98 on Artificial Analysis), and on Cursor's coding benchmark Opus at High effort beats Sonnet at Max for about 40% of the cost. The 'Sonnet beats the flagship' headline is one benchmark, run at different effort levels, inside the margin of error. Anthropic's own verdict: 'Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.'

Price and specs, side by side
| Sonnet 5.5 | Opus 5.5 | |
|---|---|---|
| Released | September 28, 2026 | September 22, 2026 |
| Input / output per 1M tokens | $2 / $10 | $4 / $20 |
| Cache read | $0.20 | $0.20 — the same, so Opus is less than 2× on cache-heavy agent loops |
| Batch | $1 / $5 | $2 / $10 |
| Fast mode | Not offered | $8 / $40 |
| Context / max output | 1M / 128K, no long-context surcharge | 1M / 128K, no long-context surcharge |
| Output speed (Artificial Analysis, max effort) | 137.6 tokens/s | 92.5 tokens/s |
| Default effort | High on the API; Medium in Claude apps and Claude Code | Medium everywhere |
| Anthropic's positioning | 'Well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets' | 'Complex work requiring careful judgment' |
The headline number, with its asterisk
Anthropic's launch table has Sonnet 5.5 at 70.6% on Terminal-Bench 4.0 versus Opus 5.5's 66.4% — the cheaper model ahead of the flagship. Three things the table's footnotes and the system card add. Sonnet was run at max effort; Opus is reported at xhigh (its max-effort score is 64.8%). The standard errors are about ±2.5 points each, so a 4.2-point gap overlaps. And safeguard fallbacks interrupted 10% of Opus's trials versus 1.5% of Sonnet's. Artificial Analysis's independent run has it 64% to 60% — same direction, same caveats. On every other row in Anthropic's own table, Opus wins: CursorBench 57.8% to 55.5%, FrontierCode 54.4% to 52.1%, Humanity's Last Exam 67.7% to 64.5%, OSWorld 81.8% to 80.1%, and a dead heat on the office-work benchmarks (GDPval 1846 to 1844). Anthropic's phrasing is the accurate one: at max effort Sonnet 'performs comparably.'
The effort dial is the whole decision
Both models have an effort setting — low, medium, high, extra high, max — that trades thinking tokens for quality. Cursor publishes score and cost per task at every level, and it's the most useful table anyone has released for this choice:
| Effort | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Low | 35.8% · $0.50 | 43.7% · $1.17 |
| Medium | 39.2% · $0.70 | 52.5% · $2.91 |
| High | 47.8% · $1.67 | 56.0% · $3.97 |
| Extra High | 53.1% · $3.88 | 56.0% · $6.98 |
| Max | 55.5% · $9.67 | 57.8% · $13.43 |
- Opus at High (56.0%, $3.97) beats Sonnet at Max (55.5%, $9.67) — better score, 41% of the cost.
- Opus at Medium (52.5%, $2.91) roughly ties Sonnet at Extra High (53.1%, $3.88) — and is cheaper.
- Sonnet's sweet spot is High and below: 47.8% for $1.67 is the best value on the table if that quality is enough for the job.
- Anthropic's developer guide says the same thing in words: Sonnet 5.5 'complements Opus 5.5 best when running at lower effort settings… At higher settings, it can perform comparably at a similar cost.' And: 'If you're tempted to use xhigh or max effort, keep in mind that Sonnet 5.5 will think longer and cost more.'
The token bill: why 'up to 30% cheaper' needs context
Anthropic's claim is that Sonnet 5.5 'costs up to 30% less per task than its predecessor' — that's versus Sonnet 5, at typical settings, and launch partners do report fewer tokens at default effort. At max effort the picture inverts. Artificial Analysis measured about 193,000 output tokens per task — 'the highest token use we have measured,' around 60% more than Opus 5.5 at max and roughly seven times GPT-6 Astra — for a cost per task of $7.60, which it puts about 50% above Sonnet 5 and above Opus 5.5's $5.98. Its summary: Sonnet 5.5 at max 'sits off the Intelligence vs. Cost per Task Pareto Frontier.' Simon Willison's first test is the same lesson at small scale: a max-effort request spent its entire 128,000-token budget thinking and produced nothing, for $1.28; the same request at extra-high cost under six cents.
Where each one wins
| If your work is… | Pick | Why |
|---|---|---|
| Bug fixes, code generation from a clear spec, data analysis | Sonnet 5.5 (medium–high) | Anthropic: it 'fits best when the task has a clear spec and a way to check the result.' Half the price, ~50% faster output |
| Documents, slides, spreadsheets, everyday writing | Sonnet 5.5 | A dead heat with Opus on the office-work benchmarks; Every's Dan Shipper reports 'dramatically improved writing even versus Opus 5.5' |
| Long autonomous coding runs, large refactors, computer use | Opus 5.5 | Anthropic's model guide: 'Complex agentic coding… multihour autonomous coding agents, large-scale refactoring' |
| Open-ended problems where you can't specify the answer | Opus 5.5 | 'Clearly stronger at complex, open-ended work requiring sustained judgment' |
| Code review where false alarms are expensive | Opus 5.5 | In CodeRabbit's test Opus caught 8 of 13 known issues at 66.7% precision; Sonnet caught 6 at 41.2% — though Sonnet cost about $0.46 a review versus $2.32 |
| Factual recall | Opus 5.5 | Artificial Analysis: 66% vs 54% accuracy on its knowledge test — though Sonnet hallucinated less often (47% vs 59%) |
| High-volume, latency-sensitive API work | Sonnet 5.5 | 137.6 vs 92.5 tokens/s; batch at $1/$5 |
If you're a Claude subscriber, not a developer
- In the Claude app: both run at medium effort by default. For everyday questions, drafting, and documents, Sonnet 5.5 is the right default and will stretch your usage further — Anthropic's support pages note that Opus costs several times more per turn than Sonnet and recommend Sonnet for most coding work. Switch to Opus for the hard one.
- Usage limits: Anthropic publishes no multiplier, only that higher effort means 'you'll reach your usage limits faster.' Practical rule: Sonnet for volume, Opus for the few tasks that matter, and neither at max unless you've hit a wall.
- Free tier: Simon Willison reports Sonnet 5.5 is 'now the model used for the free tier on claude.ai.' Anthropic's pricing page just says 'Sonnet' and the launch post doesn't address it, so treat that as reported, not confirmed.
- Claude Code: the default is still Opus 5.5 on every paid plan; version 2.1.284 added Sonnet 5.5 as the default Sonnet. The setting worth knowing is the opusplan model alias, which uses Opus to plan and Sonnet to execute — the combination Anthropic's own docs keep pointing toward.
- GitHub Copilot has Sonnet 5.5 on Pro and up (Opus 5.5 needs Pro+), billed at list price; Cursor lists both.
Sonnet 5.5 vs GPT-6 Sol — same price, different answers
We wrote last week that if Anthropic brought its pricing logic down the stack, 'the $2/$10 tier becomes a fight too.' It did, six days later, at exactly Sol's price. On quality it isn't close: Artificial Analysis has Sonnet 5.5 at 56 on its Intelligence Index to Sol's 48, and Vals has it second overall at 69.2% to Sol's 62.6%. On cost it's the reverse — $7.60 per task to Sol's $1.05 at max, $20.80 per Vals test to $7.56. The fair comparison is Sonnet at high effort, where Artificial Analysis finds it 'very narrowly behind GPT-6 Sol on Intelligence at effectively the same Cost per Task.' Same sticker, same value at matched effort; Sonnet simply has a much higher ceiling if you're willing to pay for it.
One safeguard difference worth knowing
Opus 5.5 routes biology questions to Opus 5 unless you're in Anthropic's Life Sciences Verification Program. Sonnet 5.5 doesn't — 'its biology safeguards are the same as Sonnet 5's,' with no fallback model, so life-science researchers locked out of Opus 5.5 may find Sonnet 5.5 is the most capable Claude they can actually use. Higher-risk cybersecurity tasks on Sonnet 5.5 'will visibly fall back to Sonnet 5'; it's the first Sonnet to launch with cyber safeguards.
The decision rule
- Start on Sonnet 5.5 at medium or high. It's the better default for most work and half the price.
- If the result isn't good enough, don't raise Sonnet's effort — change models. Opus at high is cheaper and better than Sonnet at max.
- Use Opus 5.5 from the start for anything long, open-ended, or expensive to get wrong.
- Never run either at max by default. It's a last resort, and on Sonnet it's the most expensive way to get a worse answer than Opus at high.
- Building agents? Opus plans, Sonnet executes.
Our read
Anthropic has shipped two models in a week that overlap almost entirely, and told you honestly which is which. The interesting design fact is that they're not tiers of intelligence so much as tiers of appetite: Sonnet gets to Opus-level scores by thinking three times as long, and Opus gets there by being smarter per token. That makes the old rule — cheaper model for easy things, expensive model for hard things — slightly wrong. The new rule is about effort, not name. For subscribers it's a quiet upgrade to the everyday Claude; for developers it's a reason to log cost per task instead of cost per token. We've updated our Claude review and Claude Code review with Sonnet 5.5, and we'll revisit once Haiku 5.5 completes the family.
