GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash: The September 2026 Benchmark Round-Up
The head-to-head
September 2026 was the busiest model-launch month in memory. OpenAI shipped GPT-6 Astra on September 3, Anthropic released Claude Fable 5.1 on September 1, and Google followed with Gemini 3.8 Flash on September 2. All three launched within a 48-hour window, and every vendor has a different benchmark number for the same test. Here is the consolidated benchmark comparison across the five benchmarks that matter for production builders.
The scores that matter
| Benchmark | GPT-6 Astra | Claude Fable 5.1 | Gemini 3.8 Flash |
|---|---|---|---|
| FrontierMath Tier 4 | 97.6% (medium effort) | 87.8% (max effort) | — |
| Artificial Analysis Index | 61 (max) | 66 (adaptive reasoning, max) | 59 (high reasoning) |
| ARC-AGI-3 (official) | 62.7% | — | — |
| ARC-AGI-3 (with adapter) | 99.9% | — | — |
| DeepSWE v1.1 | Match leader | Slightly behind | Behind |
| Terminal-Bench Science 0.1 | Leader | Trailing | Behind |
What the numbers actually say
GPT-6 Astra leads on raw capability scores. Its 97.6% on FrontierMath Tier 4 is the highest score anyone has recorded on that benchmark — though it was achieved at medium effort, not max. OpenAI also claims 99.9% on ARC-AGI-3 when run through a provider adapter, a figure that ARC Prize has not independently verified but which matches the pattern of vendor-harness scores being higher than independent evals. Astra's cost: $10/$50 per 1M tokens. Our full AGI explainer breaks down whether Astra counts as AGI.
Claude Fable 5.1 wins the Artificial Analysis Intelligence Index at 66 versus Astra's 61. Fable 5.1 achieved its score using adaptive reasoning at max effort, which is more compute-intensive than Astra's medium-effort run. The gap between the two on that index is within noise — but Fable 5.1's edge shows up more consistently across agentic coding tasks. It costs $10/$50 per 1M tokens, matching Astra's sticker price.
Gemini 3.8 Flash priced at $0.75/$3.75 per 1M tokens through December 31, 2026, then doubling to $1.50/$7.50. Its AA Index score of 59 at high reasoning is solidly competitive and at a fraction of the cost-per-task. Google's Flash tier is designed for volume workloads, not frontier scores. Our Gemini 4 coverage tracks the broader Gemini roadmap.
Cost-per-task: the real differentiator
Independent analysis by Artificial Analysis breaks down the cost-per-task across representative workloads:
| Model | Cost per benchmark task | Notes |
|---|---|---|
| GPT-6 Astra (max) | $7.63 | Highest raw score, highest cost |
| Claude Fable 5.1 (max) | $7.63 | Same cost-per-task as Astra |
| Gemini 3.8 Flash (high) | $1.24 | 6x cheaper per task than the frontier models |
The pricing gap narrows when you account for the fact that Astra and Fable 5.1 achieve more per call — fewer iterations to solve the same task. But for high-volume production use, Gemini 3.8 Flash's cost-per-task advantage is decisive.
The agent gap
All three models now support agentic tool use natively. Astra's agent harness, introduced alongside the model on September 3, ties into the OpenAI Agent SDK — but OpenAI also launched Dots, its always-on agent layer, which is the production surface most users will hit first. Fable 5.1's computer-use capabilities lead on real-browser tasks. Gemini 3.8 Flash's agent mode is competitive but trails on multi-step reasoning chains.
Bottom line for builders
| Use case | Pick |
|---|---|
| Maximum raw capability | GPT-6 Astra |
| Best balanced score/cost | Claude Fable 5.1 |
| Highest throughput, lowest cost | Gemini 3.8 Flash |
| Agentic coding | Claude Fable 5.1 (edge) or GPT-6 Astra (close) |
This benchmark comparison shows the frontier compressing. Three months ago, only GPT-5.6 Sol and Claude Opus 5 traded blows at the top. Now Gemini Flash is within striking distance of the leaders at a fraction of the cost, and OpenAI has already shipped a follow-up (GPT-6.1 Sol) that matches Astra at one-fifth the price. The benchmark round-up that matters is the one happening in production — where your actual workload, not a vendor's chosen benchmark, is the only number that counts.
Related AIPress coverage
- GPT-6 Astra Launches: OpenAI Ships 'AGI-Level' Model as Anthropic's CEO Demands a Slowdown
- What Is AGI? Artificial General Intelligence Explained Simply
- Jensen Huang Declares AGI Has Arrived: The Nvidia CEO's GPT-6 Astra Call and What's Behind It
- Claude Fable 5.1 Review: Leads Opus 5 on Every Benchmark — And Costs Less
- Gemini 4: Google DeepMind's New Chief Wants It Out 'As Soon As Possible'
- OpenAI Dots: Always-On AI Agents Built to Handle Everything
Jacob Bloom is the editor and lead writer of AIPress, covering AI model launches, benchmarks, and AI safety. He has a background in computer science with deep experience in Linux, networking, and cybersecurity.