Top 5 in AI

Signals

Gemini 4 Argon Is Real: Who Can Use It, What It Costs, and How It Actually Compares to Opus 5.5 and GPT-6 Astra

By the Top5Apps editorial team · Published September 30, 2026 · Updated September 30, 2026 · 5 min read

Share

Short answer: Gemini 4 exists, it's called Gemini 4 Argon, and almost nobody can use it yet. Google announced it on September 30 — Koray Kavukcuoglu's byline, seven days after he said it was in early post-training — as 'our new frontier model… rolling out to a set of trusted cyber defenders through our Fairwind Program,' then 'to developers, enterprises, and consumers as soon as possible,' 'starting with paid API customers and Google AI Ultra subscribers.' Not the Gemini app's free tier, not the $19.99 AI Pro plan, and no date for either. Pricing: $2 in / $10 out per million tokens as an introductory rate, rising to $4/$20. On quality, three independent scoreboards published within half an hour and disagreed: Vals ranks it #1, Arena ranks it #1 for text, and Artificial Analysis scores it 53 — tied with GPT-6 Astra and below Claude Opus 5.5 (58) and Sonnet 5.5 (56). The leaked chart we fact-checked on Sunday was wrong on every number that can now be compared.

Who can actually use it

Status
Cyber defenders in Google's Fairwind ProgramNow — 'without cyber guardrails,' usable inside CodeMender; the program has 'over 650 partners'
Paid API customers and Google AI Ultra ($99.99+)Next, 'as soon as possible' — no date; no model ID on Google's API pricing or models pages as of this evening
Developers, enterprises, consumers broadlyAfter that; Gemini Enterprise access 'supports zero data retention' per the Fairwind FAQ
Gemini app free tier and AI ProNot mentioned anywhere. The app still runs the 3.8 generation
GitHub Copilot, Cursor, OpenRouter, Vercel AI GatewayNone list it yet
From Google's announcement and Fairwind Program page, September 30, 2026. Google is also 'actively engaged in the U.S. government's voluntary process for pre-release model access.'

Google's product lead Tulsee Doshi told CNBC the staged rollout 'gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible.' Sundar Pichai's framing was blunter about why today: 'Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible.' Read that as what it is — a response to a week of leaks and to Anthropic and OpenAI both shipping — and read the rollout as what it is: an announcement, with availability to follow.

The real prices

ModelInput / output per 1M tokensNotes
Gemini 4 Argon (introductory)$2 / $10Cached input 95% off; 'after the introductory period expires' it's $4 / $20. No end date given
Gemini 4 Argon (standard)$4 / $20Same as Opus 5.5
Claude Opus 5.5$4 / $20Cached $0.20
Claude Sonnet 5.5$2 / $10Cached $0.20
GPT-6 Astra$10 / $50OpenAI's flagship
GPT-6.1 Sol$2 / $10Launched at DevDay Sept 29
Gemini 3.8 Flash$0.75 / $3.75 through Dec 31Google's current workhorse
Google's announcement and pricing page; Anthropic and OpenAI pricing pages. Google says Argon's output limit is 'an industry-leading 1M tokens, up from the previous 64K' — Vals lists 262K per call, and Artificial Analysis describes a new 'Long Decode Continuation' feature that resumes long outputs across calls, which isn't in Google's docs yet.

Google's table, and its footnotes

Pichai's chart compares Argon with GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5 across 18 benchmarks. Argon leads or ties on 13. The headline wins: Vals Index 68.9 vs 67.0 for Opus 5.5; AutomationBench 51.3 vs 42.5; DeepSWE v1.1 77.9 vs 74.2; a long-context graph test at 84.2 vs 66.8; LVBench video 91.7 vs 83.7; Harvey's legal benchmark 19.6 vs 3.8. The five it trails are the coding ones people care most about: Terminal-Bench 4.0 (57.4 vs Opus 5.5's 66.4), FrontierSWE v2 (55.0 vs Astra's 65.5), PostTrainBench (45.3 vs 49.3), Terminal-Bench Science (57.6 vs Astra's 68.1), and OSWorld 2.0 (69.2 vs Astra's 72.6). Google's methodology note is candid that rivals' numbers 'are sourced from providers' self reported numbers' and Argon's are 'self computed' — the same asymmetry every lab's launch table has, and the reason to wait for the independents.

What the independents said in the first hour

ScoreboardGemini 4 ArgonWhat it means
Artificial Analysis Intelligence Index53 — #8 of 223; 'matching GPT-6 Astra (max, 53)'Below Opus 5.5 (58) and Sonnet 5.5 (56); 23 points above Gemini 3.1 Pro. At $1.99 per task, '60% of the Cost per Task' of Astra
Vals Index#1 of 41 at 68.90%, $15.68 per testAhead of Sonnet 5.5 (67.04%) and Opus 5.5 (66.97%); #1 on its finance-agent benchmark; #5 on Terminal-Bench 4.0
Arena (text)#1, 1525, 'Preliminary,' ~4,900 votes20 points clear of the next model; #8 in the WebDev code arena
Artificial Analysis, coding detailTerminal-Bench 4.0: 57%'Only behind Claude Sonnet 5.5 (max, 64%), Claude Opus 5.5 (max, 60%) and GPT-6 Astra (59%)'
Hallucination (AA Omniscience)15% hallucination rateVersus 51% for Astra — the strongest independent number in Argon's favor
All figures as posted September 30, within about 30 minutes of the announcement; Arena marks its score preliminary. Artificial Analysis had not yet published a Coding Agent Index result.

Three boards, three answers, and they're measuring different things. Vals leans on agentic office and finance work, where Argon is strongest; Arena is human preference in chat, where a new Google model with a good voice does well; Artificial Analysis blends reasoning and coding, where Argon lands exactly where Google's own table put it — frontier on knowledge work and long context, a step behind Anthropic on terminal coding. Bloomberg's same-day report of employee skepticism about coding ('struggles to handle certain coding tasks,' which Google called 'inaccurate') fits the evidence rather than contradicting it. The fair one-line summary: Argon matches GPT-6 Astra at 60% of the cost, and Opus 5.5 remains the coding model to beat.

How the leaked chart held up

Leaked 'Gemini 4 Pro' claim (Sept 27)Actual Gemini 4 Argon (Sept 30)
DeepSWE v1.1: 88.7%77.9%
Terminal-Bench 2.1: 95.3%Google published Terminal-Bench 4.0: 57.4% — and no 2.1 number
OSWorld 2.0: 86.8%69.2% (offline subset, partial score)
'HLE-Verified': 72.1%Not in Google's table
2M-token contextThird parties list 1M; Google's post doesn't say
Prices below rivals$2/$10 introductory, then $4/$20 — Opus 5.5's price
Codename 'Argon'Correct — the one thing the leakers got right
The 'Argon' checkpoint rumored on Arena under a Flash label was never confirmed by Arena, which shows no earlier entry.

The safety posture, and the model card that isn't there

Google is leading with cyber: Argon goes to defenders 'without cyber guardrails,' and its post claims the model is 'our most resilient model yet against indirect prompt injections' — 0.7% attack success on Gray Swan's test, versus 1.0% for Opus 5.5 and 8.5% for GPT-6 Astra, by Google's chart. It also says it is 'deploying misalignment mitigations that monitor Argon's chain-of-thought and actions and stop execution when necessary,' that it used a similar monitor on training runs with alerts to 'a dedicated incident response team,' and that it's 'hardening our sandboxed environments by isolating and sealing them before high-risk training or evaluations begin' — language that reads as a direct answer to the month's agent incidents across OpenAI, Anthropic, Meta, and Google itself. What's missing: a model card (the URL 404s), a knowledge cutoff, a confirmed model ID, and any statement on the May evaluation breakout Google confirmed on September 18.

What it changes for you

  • Gemini app user, free or AI Pro: nothing today. You're on the 3.8 generation and Google has said nothing about when — or whether at your tier — that changes. Our Gemini review stands.
  • AI Ultra subscriber or API developer: you're next, at $2/$10 for an introductory period. If your work is long-context, document-heavy, or agentic office tasks, the independent numbers say Argon is worth testing against Sonnet 5.5 at the same price. If it's terminal coding, Opus 5.5 still leads every board.
  • Enterprise: Argon reaches Gemini Enterprise with zero data retention per Google's FAQ — timing unstated.
  • Everyone: the $4/$20 standard price is the tell. Google is pricing its flagship at Anthropic's flagship price, not above it — and launching at half that.

Our read

Google shipped the announcement before the product, and said so. Pichai's 'early look as soon as possible' is honest about the motive: a week of fake charts, two rivals' launches, and a model that was ready enough to benchmark but not to hand out. What's real is good — a Google flagship that matches OpenAI's at 60% of the cost, leads on long context, hallucinates far less, and comes with the strongest anti-injection numbers any lab has published. What's not there is what a reader can use: no app access, no API row, no card. We updated Sunday's fact-check the moment this dropped, because it was right — Gemini 4 wasn't out, and the chart was fiction — and because the day it becomes wrong is the day Argon shows up in a Gemini app plan. We'll re-score the Gemini review then.