Signals
Gemini 4 Argon Is Real: Who Can Use It, What It Costs, and How It Actually Compares to Opus 5.5 and GPT-6 Astra
By the Top5Apps editorial team · Published September 30, 2026 · Updated September 30, 2026 · 5 min read
Short answer: Gemini 4 exists, it's called Gemini 4 Argon, and almost nobody can use it yet. Google announced it on September 30 — Koray Kavukcuoglu's byline, seven days after he said it was in early post-training — as 'our new frontier model… rolling out to a set of trusted cyber defenders through our Fairwind Program,' then 'to developers, enterprises, and consumers as soon as possible,' 'starting with paid API customers and Google AI Ultra subscribers.' Not the Gemini app's free tier, not the $19.99 AI Pro plan, and no date for either. Pricing: $2 in / $10 out per million tokens as an introductory rate, rising to $4/$20. On quality, three independent scoreboards published within half an hour and disagreed: Vals ranks it #1, Arena ranks it #1 for text, and Artificial Analysis scores it 53 — tied with GPT-6 Astra and below Claude Opus 5.5 (58) and Sonnet 5.5 (56). The leaked chart we fact-checked on Sunday was wrong on every number that can now be compared.
Who can actually use it
| Status | |
|---|---|
| Cyber defenders in Google's Fairwind Program | Now — 'without cyber guardrails,' usable inside CodeMender; the program has 'over 650 partners' |
| Paid API customers and Google AI Ultra ($99.99+) | Next, 'as soon as possible' — no date; no model ID on Google's API pricing or models pages as of this evening |
| Developers, enterprises, consumers broadly | After that; Gemini Enterprise access 'supports zero data retention' per the Fairwind FAQ |
| Gemini app free tier and AI Pro | Not mentioned anywhere. The app still runs the 3.8 generation |
| GitHub Copilot, Cursor, OpenRouter, Vercel AI Gateway | None list it yet |
Google's product lead Tulsee Doshi told CNBC the staged rollout 'gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible.' Sundar Pichai's framing was blunter about why today: 'Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible.' Read that as what it is — a response to a week of leaks and to Anthropic and OpenAI both shipping — and read the rollout as what it is: an announcement, with availability to follow.
The real prices
| Model | Input / output per 1M tokens | Notes |
|---|---|---|
| Gemini 4 Argon (introductory) | $2 / $10 | Cached input 95% off; 'after the introductory period expires' it's $4 / $20. No end date given |
| Gemini 4 Argon (standard) | $4 / $20 | Same as Opus 5.5 |
| Claude Opus 5.5 | $4 / $20 | Cached $0.20 |
| Claude Sonnet 5.5 | $2 / $10 | Cached $0.20 |
| GPT-6 Astra | $10 / $50 | OpenAI's flagship |
| GPT-6.1 Sol | $2 / $10 | Launched at DevDay Sept 29 |
| Gemini 3.8 Flash | $0.75 / $3.75 through Dec 31 | Google's current workhorse |
Google's table, and its footnotes
Pichai's chart compares Argon with GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5 across 18 benchmarks. Argon leads or ties on 13. The headline wins: Vals Index 68.9 vs 67.0 for Opus 5.5; AutomationBench 51.3 vs 42.5; DeepSWE v1.1 77.9 vs 74.2; a long-context graph test at 84.2 vs 66.8; LVBench video 91.7 vs 83.7; Harvey's legal benchmark 19.6 vs 3.8. The five it trails are the coding ones people care most about: Terminal-Bench 4.0 (57.4 vs Opus 5.5's 66.4), FrontierSWE v2 (55.0 vs Astra's 65.5), PostTrainBench (45.3 vs 49.3), Terminal-Bench Science (57.6 vs Astra's 68.1), and OSWorld 2.0 (69.2 vs Astra's 72.6). Google's methodology note is candid that rivals' numbers 'are sourced from providers' self reported numbers' and Argon's are 'self computed' — the same asymmetry every lab's launch table has, and the reason to wait for the independents.
What the independents said in the first hour
| Scoreboard | Gemini 4 Argon | What it means |
|---|---|---|
| Artificial Analysis Intelligence Index | 53 — #8 of 223; 'matching GPT-6 Astra (max, 53)' | Below Opus 5.5 (58) and Sonnet 5.5 (56); 23 points above Gemini 3.1 Pro. At $1.99 per task, '60% of the Cost per Task' of Astra |
| Vals Index | #1 of 41 at 68.90%, $15.68 per test | Ahead of Sonnet 5.5 (67.04%) and Opus 5.5 (66.97%); #1 on its finance-agent benchmark; #5 on Terminal-Bench 4.0 |
| Arena (text) | #1, 1525, 'Preliminary,' ~4,900 votes | 20 points clear of the next model; #8 in the WebDev code arena |
| Artificial Analysis, coding detail | Terminal-Bench 4.0: 57% | 'Only behind Claude Sonnet 5.5 (max, 64%), Claude Opus 5.5 (max, 60%) and GPT-6 Astra (59%)' |
| Hallucination (AA Omniscience) | 15% hallucination rate | Versus 51% for Astra — the strongest independent number in Argon's favor |
Three boards, three answers, and they're measuring different things. Vals leans on agentic office and finance work, where Argon is strongest; Arena is human preference in chat, where a new Google model with a good voice does well; Artificial Analysis blends reasoning and coding, where Argon lands exactly where Google's own table put it — frontier on knowledge work and long context, a step behind Anthropic on terminal coding. Bloomberg's same-day report of employee skepticism about coding ('struggles to handle certain coding tasks,' which Google called 'inaccurate') fits the evidence rather than contradicting it. The fair one-line summary: Argon matches GPT-6 Astra at 60% of the cost, and Opus 5.5 remains the coding model to beat.
How the leaked chart held up
| Leaked 'Gemini 4 Pro' claim (Sept 27) | Actual Gemini 4 Argon (Sept 30) |
|---|---|
| DeepSWE v1.1: 88.7% | 77.9% |
| Terminal-Bench 2.1: 95.3% | Google published Terminal-Bench 4.0: 57.4% — and no 2.1 number |
| OSWorld 2.0: 86.8% | 69.2% (offline subset, partial score) |
| 'HLE-Verified': 72.1% | Not in Google's table |
| 2M-token context | Third parties list 1M; Google's post doesn't say |
| Prices below rivals | $2/$10 introductory, then $4/$20 — Opus 5.5's price |
| Codename 'Argon' | Correct — the one thing the leakers got right |
The safety posture, and the model card that isn't there
Google is leading with cyber: Argon goes to defenders 'without cyber guardrails,' and its post claims the model is 'our most resilient model yet against indirect prompt injections' — 0.7% attack success on Gray Swan's test, versus 1.0% for Opus 5.5 and 8.5% for GPT-6 Astra, by Google's chart. It also says it is 'deploying misalignment mitigations that monitor Argon's chain-of-thought and actions and stop execution when necessary,' that it used a similar monitor on training runs with alerts to 'a dedicated incident response team,' and that it's 'hardening our sandboxed environments by isolating and sealing them before high-risk training or evaluations begin' — language that reads as a direct answer to the month's agent incidents across OpenAI, Anthropic, Meta, and Google itself. What's missing: a model card (the URL 404s), a knowledge cutoff, a confirmed model ID, and any statement on the May evaluation breakout Google confirmed on September 18.
What it changes for you
- Gemini app user, free or AI Pro: nothing today. You're on the 3.8 generation and Google has said nothing about when — or whether at your tier — that changes. Our Gemini review stands.
- AI Ultra subscriber or API developer: you're next, at $2/$10 for an introductory period. If your work is long-context, document-heavy, or agentic office tasks, the independent numbers say Argon is worth testing against Sonnet 5.5 at the same price. If it's terminal coding, Opus 5.5 still leads every board.
- Enterprise: Argon reaches Gemini Enterprise with zero data retention per Google's FAQ — timing unstated.
- Everyone: the $4/$20 standard price is the tell. Google is pricing its flagship at Anthropic's flagship price, not above it — and launching at half that.
Our read
Google shipped the announcement before the product, and said so. Pichai's 'early look as soon as possible' is honest about the motive: a week of fake charts, two rivals' launches, and a model that was ready enough to benchmark but not to hand out. What's real is good — a Google flagship that matches OpenAI's at 60% of the cost, leads on long context, hallucinates far less, and comes with the strongest anti-injection numbers any lab has published. What's not there is what a reader can use: no app access, no API row, no card. We updated Sunday's fact-check the moment this dropped, because it was right — Gemini 4 wasn't out, and the chart was fiction — and because the day it becomes wrong is the day Argon shows up in a Gemini app plan. We'll re-score the Gemini review then.
