AI Research

GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash: The September 2026 Benchmark Round-Up

Hero image for September 2026 benchmark round-up: dark navy gradient with 'FRONTIER BENCHMARK ROUND-UP' in large type, GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash comparison, AIPress mark, abstract geometric circuit motifs

The head-to-head

September 2026 was the busiest model-launch month in memory. OpenAI shipped GPT-6 Astra on September 3, Anthropic released Claude Fable 5.1 on September 1, and Google followed with Gemini 3.8 Flash on September 2. All three launched within a 48-hour window, and every vendor has a different benchmark number for the same test. Here is the consolidated benchmark comparison across the five benchmarks that matter for production builders.

The scores that matter

Benchmark GPT-6 Astra Claude Fable 5.1 Gemini 3.8 Flash
FrontierMath Tier 4 97.6% (medium effort) 87.8% (max effort) —
Artificial Analysis Index 61 (max) 66 (adaptive reasoning, max) 59 (high reasoning)
ARC-AGI-3 (official) 62.7% — —
ARC-AGI-3 (with adapter) 99.9% — —
DeepSWE v1.1 Match leader Slightly behind Behind
Terminal-Bench Science 0.1 Leader Trailing Behind

What the numbers actually say

GPT-6 Astra leads on raw capability scores. Its 97.6% on FrontierMath Tier 4 is the highest score anyone has recorded on that benchmark — though it was achieved at medium effort, not max. OpenAI also claims 99.9% on ARC-AGI-3 when run through a provider adapter, a figure that ARC Prize has not independently verified but which matches the pattern of vendor-harness scores being higher than independent evals. Astra's cost: $10/$50 per 1M tokens. Our full AGI explainer breaks down whether Astra counts as AGI.

Claude Fable 5.1 wins the Artificial Analysis Intelligence Index at 66 versus Astra's 61. Fable 5.1 achieved its score using adaptive reasoning at max effort, which is more compute-intensive than Astra's medium-effort run. The gap between the two on that index is within noise — but Fable 5.1's edge shows up more consistently across agentic coding tasks. It costs $10/$50 per 1M tokens, matching Astra's sticker price.

Gemini 3.8 Flash priced at $0.75/$3.75 per 1M tokens through December 31, 2026, then doubling to $1.50/$7.50. Its AA Index score of 59 at high reasoning is solidly competitive and at a fraction of the cost-per-task. Google's Flash tier is designed for volume workloads, not frontier scores. Our Gemini 4 coverage tracks the broader Gemini roadmap.

Cost-per-task: the real differentiator

Independent analysis by Artificial Analysis breaks down the cost-per-task across representative workloads:

Model Cost per benchmark task Notes
GPT-6 Astra (max) $7.63 Highest raw score, highest cost
Claude Fable 5.1 (max) $7.63 Same cost-per-task as Astra
Gemini 3.8 Flash (high) $1.24 6x cheaper per task than the frontier models

The pricing gap narrows when you account for the fact that Astra and Fable 5.1 achieve more per call — fewer iterations to solve the same task. But for high-volume production use, Gemini 3.8 Flash's cost-per-task advantage is decisive.

The agent gap

All three models now support agentic tool use natively. Astra's agent harness, introduced alongside the model on September 3, ties into the OpenAI Agent SDK — but OpenAI also launched Dots, its always-on agent layer, which is the production surface most users will hit first. Fable 5.1's computer-use capabilities lead on real-browser tasks. Gemini 3.8 Flash's agent mode is competitive but trails on multi-step reasoning chains.

Bottom line for builders

Use case Pick
Maximum raw capability GPT-6 Astra
Best balanced score/cost Claude Fable 5.1
Highest throughput, lowest cost Gemini 3.8 Flash
Agentic coding Claude Fable 5.1 (edge) or GPT-6 Astra (close)

This benchmark comparison shows the frontier compressing. Three months ago, only GPT-5.6 Sol and Claude Opus 5 traded blows at the top. Now Gemini Flash is within striking distance of the leaders at a fraction of the cost, and OpenAI has already shipped a follow-up (GPT-6.1 Sol) that matches Astra at one-fifth the price. The benchmark round-up that matters is the one happening in production — where your actual workload, not a vendor's chosen benchmark, is the only number that counts.

Related AIPress coverage


Jacob Bloom is the editor and lead writer of AIPress, covering AI model launches, benchmarks, and AI safety. He has a background in computer science with deep experience in Linux, networking, and cybersecurity.

Building something with AI?

DevsIsle designs and ships AI systems, agents and integrations for teams that need it done properly.

Talk to our team →