Top 5 in AI

Guides

What Is Claude Haiku 5.5? The Price (and What '75% Less' Actually Means), vs Haiku 4.5 and GPT-6 Luna, and How to Turn It On in Claude Code

By the Top5Apps editorial team · Published October 8, 2026 · Updated October 8, 2026 · 8 min read

Share

Short answer: Claude Haiku 5.5 is Anthropic's cheapest and fastest model, released October 7, priced at $0.10 per million input tokens and $0.50 per million output on prompts up to 100,000 tokens — 90 percent below Haiku 4.5 and identical to OpenAI's GPT-6 Luna — with a 1-million-token context window, 128K output, and an adjustable effort setting for the first time on a Haiku. Anthropic's own line is that it 'costs around 75% less to run than Claude Haiku 4.5,' and the launch post is at 5.5 million views. The 75 and the 90 are both true: 90 is the list-price cut on short prompts, 75 is Anthropic's estimate of what you'll actually save once two things are counted — a long-prompt tier above 100K tokens that costs five times more, and a new tokenizer that uses about 30 percent more tokens per task. It's in the Claude app on every plan, in Claude Code, Cursor, Copilot, Devin, Bedrock, Vertex, and Foundry today. One correction to the launch chatter: Haiku 5.5 is not Claude Code's default model or its default subagent model — Opus 5.5 remains both — so if you want the savings there, you set it, and the exact settings are below.

The price, precisely

Per million tokensInputOutputCache readBatch (in / out)
Haiku 5.5, prompts ≤100K tokens$0.10$0.50$0.01$0.05 / $0.25
Haiku 5.5, prompts >100K tokens$0.50$2.50$0.05$0.25 / $1.25
Haiku 4.5$1$5$0.10$0.50 / $2.50
GPT-6 Luna (≤272K input)$0.10$0.50$0.0150% of standard
Sonnet 5.5$2$10$0.10 (was $0.20)$1 / $5
Opus 5.5$4$20$0.20$2 / $10
From Anthropic's pricing page and OpenAI's GPT-6 Luna model page, October 8, 2026. Anthropic also halved Sonnet 5.5 cache reads to $0.10 on October 7.

Three things to read off that table. First, the cut is exactly 90 percent on every line for short prompts ($1 → $0.10, $5 → $0.50, cache reads $0.10 → $0.01) and exactly 50 percent above 100K. Anthropic's footnote explains the 75: 'Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category. This calculation also accounts for… an updated tokenizer… which means it uses slightly more tokens per task.' The docs put 'slightly' at about 30 percent. Second, Haiku 5.5 is a twentieth of Sonnet 5.5 and a fortieth of Opus 5.5 on short prompts — which is why Anthropic pitches it as the subagent next to those two rather than a replacement for them. Third, the long-prompt tier is the trap. Luna's step-up comes at 272K tokens and is 2× input / 1.5× output; Haiku's comes at 100K and is 5×. Simon Willison's read: 'Above 100,000 tokens, Luna looks like a much better deal.' If your agent loop stuffs a big context into every call, price it on the second row, not the first.

What's new besides the price

  • 1M context window and 128K output, up from 200K and 64K on Haiku 4.5 — and on the Anthropic API the full 1M comes at standard pricing on every plan, with the 100K price step the only caveat.
  • An effort setting, 'our first Haiku-class model to come with an adjustable effort setting' — low, medium (default), high, xhigh, max. Thinking is adaptive and can't be turned off at xhigh or max; manual budget_tokens now returns an error.
  • Speed: 'our fastest model to date at each model's standard speed,' with a footnote that Opus in Fast Mode is quicker. Artificial Analysis measured 242 tokens a second at max effort, 137 at medium, against about 129 for Sonnet 5.5 and Luna.
  • Knowledge cutoff June 2026; model ID claude-haiku-5-5; not retiring before October 7, 2027. Omit temperature, top_p, and top_k — any non-default value returns a 400.
  • Safeguards: more restrictive on cybersecurity than Haiku 4.5, 'somewhat less restrictive' than Sonnet 5.5 — defensive tasks allowed, penetration testing blocked; biology safeguards the same as Sonnet 5.5 and Opus 5. The system card says it 'does not cross any new RSP thresholds.'

Haiku 5.5 vs Haiku 4.5 vs GPT-6 Luna: the numbers

BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
Terminal-Bench 4.039.2%0.0%16.4%70.6%
OSWorld 2.1 (computer use)72.4%15.7%48.9%83.9%
FrontierCode 1.1, Main set46.4%—42.4%52.1%
GDPval-AA v2.1 (knowledge work, Elo)162073514371840
Humanity's Last Exam, no tools45.9%10.2%—56.9%
Chartography, no tools46.4%6.4%29.1%61.6%
Artificial Analysis Intelligence Index (independent)43173856
Anthropic's announcement table (vendor-run, Haiku 5.5 at max effort), plus the independent Artificial Analysis index read October 8. Anthropic published no SWE-bench Verified or GPQA figure; the system card adds SWE-bench Pro 64.8% vs Sonnet 5.5's 81.3%.

The shape is consistent across vendor and independent numbers: a very large jump from Haiku 4.5 (Terminal-Bench from zero to 39 percent; the AA index up 26 points in a year, which Alex Albert pointed out is 'less than a year apart'), a clear lead over Luna at the same price (43 vs 38 on AA; 39 vs 16 on terminals; 72 vs 49 on computer use), and a real gap to Sonnet 5.5 (13 points on AA, 31 on Terminal-Bench). Two independent caveats from Artificial Analysis: Haiku 5.5 at max effort 'uses ~162k output tokens per Intelligence Index task, ~3x GPT-6 Luna,' so on reasoning-heavy work the per-token price advantage over Luna narrows; and its AutomationBench score is 'likely understated' because of an over-refusal issue Anthropic says it's fixing. On Vals it debuts third of 110 on Vibe Code Bench and 16th of 45 on the overall index. Arena hadn't scored it yet when we checked.

About Cognition's 58.4%

Devin's maker posted that 'On FrontierCode 1.1, Haiku 5.5 scores 58.4%. That puts it ahead of Sonnet 5 at about an eighth of the cost per task,' and the line traveled. Two clarifications from the system card and Cognition's own chart. 58.4 is the Extended set — all 150 tasks; on the Main set of the 100 hardest, Haiku 5.5 scores 46.4, which is the number in Anthropic's table. And 'ahead of Sonnet 5' is 58.4 to 56.2 — a 2.2-point edge, with Luna at 56.1 and Kimi K3 at 58.2 in the same band; Sonnet 5.5 is 64.4 and Opus 5.5 65.3. The 'eighth of the cost' has no public dollar figures behind it. None of that makes it a bad result — a $0.50-output model within six points of Sonnet 5.5 on real engineering tasks is the whole story — but it's an Extended-set number being read as a headline one.

How to turn it on in Claude Code

Update first: Haiku 5.5 arrived in Claude Code 2.1.293 on October 7 ('Added Claude Haiku 5.5 (claude-haiku-5-5), now the default Haiku model on the Anthropic API'); run claude update. Then know what didn't change: the default model on Pro, Max, Team, Enterprise, and the API is still Opus 5.5, and subagents inherit the main conversation's model by default — Haiku 5.5 is nobody's default. The 'summaries, compactions, database queries' line in the launch is Anthropic's API pitch, not a description of Claude Code's internals; the docs tie the haiku alias to 'background functionality' generically and never say compaction runs on it. So:

  • Use it as your main model for a session: /model claude-haiku-5-5, or start with claude --model claude-haiku-5-5. The haiku alias also resolves to 5.5 — on the Anthropic API only; on Bedrock, Foundry, and the Google Agent Platform, haiku still means Haiku 4.5.
  • Make one subagent run on it: add model: haiku to that agent's frontmatter in .claude/agents/. The docs' own suggestion: define an Explore-style agent with model: haiku to 'run exploration on a lower-cost model.' Anthropic's built-in claude-code-guide helper already runs on Haiku.
  • Make every subagent run on it: set the environment variable CLAUDE_CODE_SUBAGENT_MODEL=haiku, and add CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1 to override agents that specify their own model. In settings.json: {"env": {"CLAUDE_CODE_SUBAGENT_MODEL": "haiku", "CLAUDE_CODE_SUBAGENT_MODEL_FORCE": "1"}}.
  • Point background work at it: ANTHROPIC_DEFAULT_HAIKU_MODEL is the variable for 'the haiku alias, or background functionality' (the older ANTHROPIC_SMALL_FAST_MODEL is deprecated).
  • Two behaviors to expect: 1M context on every plan with no [1m] suffix; default effort medium, and thinking can't be turned off on this model. A session saved on Haiku 4.5 resumes on 5.5 once the alias points there.
  • What you pay: on a subscription, Haiku calls draw from the same usage pool, cheaply; on the API, the table above — and Max 5x/20x and Team subscribers now get $100 / $200 / up to $500 a month in Claude Platform API credits that work 'on any model, including Haiku 5.5.'

Where else it is today

  • Claude app: selectable on Free, Pro, Max, Team, and Enterprise, web and mobile, with the effort selector. Anthropic hasn't said whether it's the Free tier's default — only that Free users 'can select' it.
  • Cursor: live the same hour ('On shorter requests, it costs 10x less than Claude Haiku 4.5'), at list price from the Other Models pool, toggled in Settings → Models; requests over 100K input bill at 5×.
  • GitHub Copilot: generally available for Pro, Pro+, Max, Business, and Enterprise (not listed for Free), 'billed at provider list pricing under usage-based billing'; GitHub's early testing 'matched Claude Sonnet 5 on many coding tasks while using significantly fewer tokens and steps.'
  • Devin: Desktop and CLI, as the 'sidekick' to an Opus 5.5 lead in Fusion — Cognition says that pairing 'holds a top-tier FrontierCode score of 66.2.' (Windsurf is now Devin Desktop.)
  • Clouds: Amazon Bedrock (US, EU, AU, JP, Global, and GovCloud profiles), Google Cloud's Agent Platform (GA, 1M in / 128K out), and Microsoft Foundry, all as claude-haiku-5-5.

Who should switch, and to what

  • Anyone running agent loops with thousands of small calls — classification, extraction, summaries, routing, tool-result triage. Prompts under 100K tokens at $0.10/$0.50 is the design target; Anthropic says 90 percent of Haiku 4.5 requests were already there.
  • Claude Code users with busy subagents: CLAUDE_CODE_SUBAGENT_MODEL=haiku is the single biggest token-bill lever this week, and Explore-type agents lose little.
  • Not for long-context work. Above 100K tokens the price quintuples and Luna is cheaper; above that, Sonnet 5.5 with its newly halved cache reads is often the better buy than Haiku on the second tier.
  • Not a Sonnet replacement on hard code. Terminal-Bench 39 vs 71 and SWE-bench Pro 65 vs 81 are the honest gaps. Use it as the sidekick, which is exactly what Anthropic, Cognition, and GitHub all say.
  • Luna users: same price, better scores on every published benchmark, 3× the tokens on hard reasoning, and a harsher long-prompt tier. Switch for agents and computer use; benchmark your own long prompts first.

Our read

This is the model that changes the bill, not the leaderboard. Haiku 4.5 was a year old and scored zero on Terminal-Bench; its successor matches GPT-6 Luna's price and beats it everywhere Anthropic and Artificial Analysis measured, while remaining a tier below Sonnet 5.5 on the work that matters most. The two numbers to carry are 90 and 100K: the cut is 90 percent until your prompt passes 100,000 tokens, after which it's a different, pricier model. The one thing to un-learn from the launch chatter is that Claude Code will do this for you — it won't; set the subagent variable and the savings are real. Our Claude review, Claude Code review, and Opus 5.5 vs Sonnet 5.5 guide are updated for the three-model lineup.

Where these apps rank