Signals
GPT-6 Astra vs Claude Fable 5.1: The Honest Scorecard
By the Top5Apps editorial team · Published September 9, 2026 · Updated September 9, 2026
The first week of September gave us the closest thing AI has had to a title fight: Anthropic shipped Claude Fable 5.1 on the 1st, OpenAI answered with GPT-6 Astra on the 3rd, and both priced their flagship at exactly $10 per million input tokens and $50 out — a list-price tie so precise it reads as a statement. We spent the week reading every benchmark table, every independent leaderboard, and every credible hands-on report. The short version: the popular take — Astra for driving software, Fable for writing it — is directionally right, but the details are where your money is. (Disclosure, per The Receipts Standard: this site is produced using Claude-family tools, which is exactly why everything below leans on third-party measurements wherever they exist.)
First: the benchmark tables are marketing
Before trusting any number from either launch post, know three things. One: the OSWorld computer-use scores circulating everywhere — Astra 72.6%, Fable 41.7% — are not the same test. Astra's number is a partial-credit score on an offline subset; Fable's is a strict score with production safeguards enabled (Anthropic's own partial-credit figure is 77.9%, on its own modified setup). Each vendor explicitly refuses to print the other's model on that row, with OpenAI's footnote accusing Anthropic's system card of 'modified tasks and modified grading.' Two: Astra's headline 99.9% on ARC-AGI-3 is real and ARC-verified — using OpenAI's own adapter harness; on ARC's standard harness the verified score is 62.71%. Both numbers are true. Only one is comparable. Three: Artificial Analysis re-versioned its Intelligence Index twice in two days after these launches — Fable 5.1 led Astra by 4.5 points on the version OpenAI's post cites, and they're now tied at 53 on the current one. When the yardsticks move this much, treat every table as an opening argument, not a verdict.
Tool use: Astra, clearly
This one isn't close, and it survives the benchmark caveats because the evidence is behavioral, not tabular. On the tests with cleanest common footing, Astra leads computer-driving decisively — 92.7% on ScreenSpot-Pro (vs 87.3% for the Fable family, and OpenAI had to substitute the safeguard-lifted Mythos variant to get Anthropic's best number), and 41.4% vs 31.4% on AutomationBench. The hands-on record agrees: Simon Willison drove a locally installed Blender through OpenAI's Codex with Astra — three prompts to a rendered, flair-added scene, about $4.24 of API time — and named practitioners have had it composing inside Ableton and turning gray-box games into themed prototypes. Claude's computer use is real and now runs in the background on desktop, but Astra was visibly built for this: it completes OSWorld-class tasks in roughly half the time of its own predecessor, and Wharton's Ethan Mollick reports it spinning up sub-agents to art-direct Blender projects unprompted. If your dream is an AI operating your creative tools, Astra is the current answer.
Coding: closer than either fanbase admits
Here's where the popular take needs surgery. On neutral-harness benchmarks, Astra actually edges ahead: Artificial Analysis's own Terminal-Bench 4.0 run puts it at 59.6% vs Fable 5.1's 55.1%, and the official Terminal-Bench leaderboard has them in a statistical tie (58.2 vs 57.9, overlapping error bars) — with Astra's run costing roughly half as much. But everywhere humans choose rather than harnesses, Fable wins: it's the highest score Cursor has ever recorded on CursorBench (73.4% — a benchmark Astra can never enter, since OpenAI barred its models from Cursor after the SpaceX acquisition), it sits #1 on arena.ai's agent leaderboard where Astra hadn't charted at last update, and Artificial Analysis's Coding Agent Index has the two functionally tied. One stat OpenAI printed against itself: on Humanity's Last Exam with tools, both vendors' tables agree Fable 5.1 beats Astra, 65.0% to 57.2%. Our read: for delegated, long-horizon engineering — the Claude Code workload — Fable keeps the crown; for fast, cheap, tool-assisted coding runs, Astra is at worst even and meaningfully cheaper per task.
The economics: same sticker, different bill
The identical $10/$50 pricing is where the real story hides. Running Artificial Analysis's full evaluation suite cost $5,324 with Astra and $13,129 with Fable 5.1 — because Fable thinks in public, emitting roughly 78,000 output tokens per task to Astra's 27,000. Same list price, ~2.4× difference in what a task actually costs. Anthropic's counterweight is caching: Fable 5.1's cache reads cost $0.25 per million against Astra's $1.00, which flips the math back for exactly the repetitive, context-heavy agent loops Claude Code runs all day (Anthropic estimates up to ~45% savings on agentic workloads versus Fable 5). Translation for buyers: Astra is cheaper per fresh task, Fable is cheaper per repeated one — and anyone quoting only the list price is telling you nothing.
The trust ledger
Both launches carried asterisks their keynotes didn't. Astra's new recurrent-depth architecture makes it dramatically more efficient — and, by OpenAI's own admission, harder to monitor: its written reasoning obscures more of what it's actually doing, a decline OpenAI's chief scientist calls fragile and trending negative. It's also OpenAI's first model rated Critical for cybersecurity, with those capabilities gated — and the rollout was messy enough that Sam Altman apologized and OpenAI issued banked-reset compensation before access stabilized on September 5 (Plus users still get Astra only in Work and Codex, not the main Chat picker). Fable 5.1's asterisks run the other direction: it shipped everywhere on day one, but it scores benchmark zeros wherever its production safeguards intervene — Anthropic tests the model you actually get, which costs it points and earns it credibility — and its output carries watermarking that some users resent. Pick your discomfort: the model that hides its reasoning, or the one that shows its restraints.
The bottom line
Refined for real buyers: Astra if AI's job is to operate your software — computer use, creative-tool automation, fast cheap agent runs — and its speed-per-dollar makes it the efficiency pick even where quality ties. Fable 5.1 if AI's job is to be your best colleague — long-horizon coding via Claude Code, writing, judgment-heavy work — with the receipts coming from where practitioners actually vote: Cursor's best-ever benchmark, the human-preference agent board, and cache economics built for sustained work. The app-level decision is a different question with a different answer — that's our ChatGPT vs Claude head-to-head — and we'll re-run this scorecard as the independent leaderboards fill in. The one take we'll defend against both fanbases: neither model 'won' September. They specialized, in public, at identical prices — and the benchmark tables were written to obscure exactly that.
