Top 5 in AI

Guides

Claude Opus 5.5 vs Sonnet 5.5: Which Should You Use? (The Cheaper One Isn't Always Cheaper)

By the Top5Apps editorial team · Published September 29, 2026 · Updated September 29, 2026 · 6 min read

Share

Short answer: use Sonnet 5.5 at medium or high effort for well-defined work, and switch to Opus 5.5 the moment you're tempted to turn Sonnet up to max. Sonnet 5.5 launched September 28 at $2 in / $10 out per million tokens — half of Opus 5.5's $4/$20 — and it's faster. But list price isn't what you pay. At max effort, Sonnet 5.5 uses so many tokens that independent testing found it costs more per task than Opus ($7.60 vs $5.98 on Artificial Analysis), and on Cursor's coding benchmark Opus at High effort beats Sonnet at Max for about 40% of the cost. The 'Sonnet beats the flagship' headline is one benchmark, run at different effort levels, inside the margin of error. Anthropic's own verdict: 'Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.'

The Claude logo — a white starburst mark beside the word Claude — over a view of Earth's blue horizon seen through a spacecraft window
Anthropic's Claude 5.5 family now has two members, Opus 5.5 and Sonnet 5.5, with Haiku 5.5 to come. Image: Anthropic.

Price and specs, side by side

Sonnet 5.5Opus 5.5
ReleasedSeptember 28, 2026September 22, 2026
Input / output per 1M tokens$2 / $10$4 / $20
Cache read$0.20$0.20 — the same, so Opus is less than 2× on cache-heavy agent loops
Batch$1 / $5$2 / $10
Fast modeNot offered$8 / $40
Context / max output1M / 128K, no long-context surcharge1M / 128K, no long-context surcharge
Output speed (Artificial Analysis, max effort)137.6 tokens/s92.5 tokens/s
Default effortHigh on the API; Medium in Claude apps and Claude CodeMedium everywhere
Anthropic's positioning'Well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets''Complex work requiring careful judgment'
From Anthropic's pricing and model pages, Sept 29, 2026. Sonnet 5 stays at $2/$10 too — a planned rise to $3/$15 'will not occur.' Haiku 5.5 'will join the Claude 5.5 family in the coming weeks.'

The headline number, with its asterisk

Anthropic's launch table has Sonnet 5.5 at 70.6% on Terminal-Bench 4.0 versus Opus 5.5's 66.4% — the cheaper model ahead of the flagship. Three things the table's footnotes and the system card add. Sonnet was run at max effort; Opus is reported at xhigh (its max-effort score is 64.8%). The standard errors are about ±2.5 points each, so a 4.2-point gap overlaps. And safeguard fallbacks interrupted 10% of Opus's trials versus 1.5% of Sonnet's. Artificial Analysis's independent run has it 64% to 60% — same direction, same caveats. On every other row in Anthropic's own table, Opus wins: CursorBench 57.8% to 55.5%, FrontierCode 54.4% to 52.1%, Humanity's Last Exam 67.7% to 64.5%, OSWorld 81.8% to 80.1%, and a dead heat on the office-work benchmarks (GDPval 1846 to 1844). Anthropic's phrasing is the accurate one: at max effort Sonnet 'performs comparably.'

The effort dial is the whole decision

Both models have an effort setting — low, medium, high, extra high, max — that trades thinking tokens for quality. Cursor publishes score and cost per task at every level, and it's the most useful table anyone has released for this choice:

EffortSonnet 5.5Opus 5.5
Low35.8% · $0.5043.7% · $1.17
Medium39.2% · $0.7052.5% · $2.91
High47.8% · $1.6756.0% · $3.97
Extra High53.1% · $3.8856.0% · $6.98
Max55.5% · $9.6757.8% · $13.43
CursorBench 4.0 score and cost per task, from cursor.com/evals. Sonnet's per-task costs are Anthropic's estimates from Cursor's token counts. Cursor is owned by SpaceX, a competitor to Anthropic; the numbers are still the only public per-effort comparison.
  • Opus at High (56.0%, $3.97) beats Sonnet at Max (55.5%, $9.67) — better score, 41% of the cost.
  • Opus at Medium (52.5%, $2.91) roughly ties Sonnet at Extra High (53.1%, $3.88) — and is cheaper.
  • Sonnet's sweet spot is High and below: 47.8% for $1.67 is the best value on the table if that quality is enough for the job.
  • Anthropic's developer guide says the same thing in words: Sonnet 5.5 'complements Opus 5.5 best when running at lower effort settings… At higher settings, it can perform comparably at a similar cost.' And: 'If you're tempted to use xhigh or max effort, keep in mind that Sonnet 5.5 will think longer and cost more.'

The token bill: why 'up to 30% cheaper' needs context

Anthropic's claim is that Sonnet 5.5 'costs up to 30% less per task than its predecessor' — that's versus Sonnet 5, at typical settings, and launch partners do report fewer tokens at default effort. At max effort the picture inverts. Artificial Analysis measured about 193,000 output tokens per task — 'the highest token use we have measured,' around 60% more than Opus 5.5 at max and roughly seven times GPT-6 Astra — for a cost per task of $7.60, which it puts about 50% above Sonnet 5 and above Opus 5.5's $5.98. Its summary: Sonnet 5.5 at max 'sits off the Intelligence vs. Cost per Task Pareto Frontier.' Simon Willison's first test is the same lesson at small scale: a max-effort request spent its entire 128,000-token budget thinking and produced nothing, for $1.28; the same request at extra-high cost under six cents.

Where each one wins

If your work is…PickWhy
Bug fixes, code generation from a clear spec, data analysisSonnet 5.5 (medium–high)Anthropic: it 'fits best when the task has a clear spec and a way to check the result.' Half the price, ~50% faster output
Documents, slides, spreadsheets, everyday writingSonnet 5.5A dead heat with Opus on the office-work benchmarks; Every's Dan Shipper reports 'dramatically improved writing even versus Opus 5.5'
Long autonomous coding runs, large refactors, computer useOpus 5.5Anthropic's model guide: 'Complex agentic coding… multihour autonomous coding agents, large-scale refactoring'
Open-ended problems where you can't specify the answerOpus 5.5'Clearly stronger at complex, open-ended work requiring sustained judgment'
Code review where false alarms are expensiveOpus 5.5In CodeRabbit's test Opus caught 8 of 13 known issues at 66.7% precision; Sonnet caught 6 at 41.2% — though Sonnet cost about $0.46 a review versus $2.32
Factual recallOpus 5.5Artificial Analysis: 66% vs 54% accuracy on its knowledge test — though Sonnet hallucinated less often (47% vs 59%)
High-volume, latency-sensitive API workSonnet 5.5137.6 vs 92.5 tokens/s; batch at $1/$5
CodeRabbit is an Anthropic launch partner. One independent bug-hunt (small samples) found Sonnet 5.5 at max the top scorer and 'the least lazy model I tested' — using roughly three times the turns of Opus to get there.

If you're a Claude subscriber, not a developer

  • In the Claude app: both run at medium effort by default. For everyday questions, drafting, and documents, Sonnet 5.5 is the right default and will stretch your usage further — Anthropic's support pages note that Opus costs several times more per turn than Sonnet and recommend Sonnet for most coding work. Switch to Opus for the hard one.
  • Usage limits: Anthropic publishes no multiplier, only that higher effort means 'you'll reach your usage limits faster.' Practical rule: Sonnet for volume, Opus for the few tasks that matter, and neither at max unless you've hit a wall.
  • Free tier: Simon Willison reports Sonnet 5.5 is 'now the model used for the free tier on claude.ai.' Anthropic's pricing page just says 'Sonnet' and the launch post doesn't address it, so treat that as reported, not confirmed.
  • Claude Code: the default is still Opus 5.5 on every paid plan; version 2.1.284 added Sonnet 5.5 as the default Sonnet. The setting worth knowing is the opusplan model alias, which uses Opus to plan and Sonnet to execute — the combination Anthropic's own docs keep pointing toward.
  • GitHub Copilot has Sonnet 5.5 on Pro and up (Opus 5.5 needs Pro+), billed at list price; Cursor lists both.

Sonnet 5.5 vs GPT-6 Sol — same price, different answers

We wrote last week that if Anthropic brought its pricing logic down the stack, 'the $2/$10 tier becomes a fight too.' It did, six days later, at exactly Sol's price. On quality it isn't close: Artificial Analysis has Sonnet 5.5 at 56 on its Intelligence Index to Sol's 48, and Vals has it second overall at 69.2% to Sol's 62.6%. On cost it's the reverse — $7.60 per task to Sol's $1.05 at max, $20.80 per Vals test to $7.56. The fair comparison is Sonnet at high effort, where Artificial Analysis finds it 'very narrowly behind GPT-6 Sol on Intelligence at effectively the same Cost per Task.' Same sticker, same value at matched effort; Sonnet simply has a much higher ceiling if you're willing to pay for it.

One safeguard difference worth knowing

Opus 5.5 routes biology questions to Opus 5 unless you're in Anthropic's Life Sciences Verification Program. Sonnet 5.5 doesn't — 'its biology safeguards are the same as Sonnet 5's,' with no fallback model, so life-science researchers locked out of Opus 5.5 may find Sonnet 5.5 is the most capable Claude they can actually use. Higher-risk cybersecurity tasks on Sonnet 5.5 'will visibly fall back to Sonnet 5'; it's the first Sonnet to launch with cyber safeguards.

The decision rule

  • Start on Sonnet 5.5 at medium or high. It's the better default for most work and half the price.
  • If the result isn't good enough, don't raise Sonnet's effort — change models. Opus at high is cheaper and better than Sonnet at max.
  • Use Opus 5.5 from the start for anything long, open-ended, or expensive to get wrong.
  • Never run either at max by default. It's a last resort, and on Sonnet it's the most expensive way to get a worse answer than Opus at high.
  • Building agents? Opus plans, Sonnet executes.

Our read

Anthropic has shipped two models in a week that overlap almost entirely, and told you honestly which is which. The interesting design fact is that they're not tiers of intelligence so much as tiers of appetite: Sonnet gets to Opus-level scores by thinking three times as long, and Opus gets there by being smarter per token. That makes the old rule — cheaper model for easy things, expensive model for hard things — slightly wrong. The new rule is about effort, not name. For subscribers it's a quiet upgrade to the everyday Claude; for developers it's a reason to log cost per task instead of cost per token. We've updated our Claude review and Claude Code review with Sonnet 5.5, and we'll revisit once Haiku 5.5 completes the family.