GPT-6 Astra Rated "Critical" on ExploitBench — the Gap Between Score and Classification
OpenAI's GPT-6 Astra — the model that scored a perfect 100% on the company's own ExploitBench cyber benchmark on September 4 — has now been rated "Critical" on the same test, according to a report published September 25 by tech-insider.org.
The shift from "perfect score" to "Critical" is not a contradiction. It is the difference between raw capability and the judgment that follows it: the model can do something dangerous, and the people measuring it have decided that what it can do crosses a line.
What ExploitBench measures
ExploitBench is OpenAI's own cybersecurity evaluation — a red-team-style test in which models are asked to find, chain, and weaponize real software vulnerabilities. It is the hardest cyber benchmark OpenAI publishes, and the one it uses to decide whether a model's offensive capabilities require tighter controls.
On September 4, GPT-6 Astra became the first model to score 100% on the test. The Hacker News covered the result the same day, alongside reporting from Yellow.com, Gagadget, and CryptoRank. OpenAI itself framed the milestone in a September 1 "Path to Astra" post as a capability checkpoint, not a safety verdict — the model had demonstrated "critical capabilities" under "frontier safeguards."
The September 25 rating changes the framing. A "Critical" label on ExploitBench is not a score; it is a risk classification. It says the model's combination of exploit-finding, exploit-chaining, and payload-generation ability is now at a level that the evaluator — in this case the same test OpenAI designed — considers serious enough to warrant a formal warning.
The timeline that matters
- September 1, 2026: OpenAI launches GPT-6 Astra commercially, alongside the "Path to Astra" capabilities post.
- September 4, 2026: GPT-6 Astra scores 100% on ExploitBench. Multiple outlets cover the milestone. OpenAI does not issue a new safety statement specifically on the score.
- September 25, 2026: tech-insider.org reports the model is now rated "Critical" on the same benchmark.
The 21-day gap between the score and the rating is the story's quiet center. A model can score perfectly on a cyber test on a Thursday and be formally classified as a critical-risk system by the following Friday — and in between those two dates, it is available to anyone with an API key.
What "Critical" does and does not mean
A Critical rating on a benchmark is not the same as a regulatory designation. There is no formal "Critical AI model" category in U.S. law today. It does not automatically trigger export controls, deployment restrictions, or mandatory disclosure — at least not under the frameworks that exist right now.
What it does is more subtle and, in some ways, more important: it is OpenAI's own test, run by OpenAI's own standards, telling OpenAI that the model it shipped three weeks earlier can now do things that OpenAI's own rubric considers critical.
That is a rare and notable kind of statement from a frontier lab. Labs usually prefer to publish capability scores without attaching a risk verdict to the same number. A lab publishing both — here is what GPT-6 Astra can do, and here is the risk level we assign to that capability — is a different model of transparency, whether it is deliberate or simply what happens when the score keeps climbing.
The context OpenAI is operating in
This GPT-6 Astra Critical rating lands inside a crowded September for OpenAI on the safety and governance front:
- On September 19, OpenAI disclosed six new AI safety incidents and launched a misalignment reporting framework — see OpenAI Discloses Six New AI Safety Incidents and Launches Misalignment Reporting Framework.
- On September 24, OpenAI, Anthropic, and Hugging Face CEOs briefed the UN Security Council on global AI regulation — the first time the Council has taken up the question on its own — see OpenAI, Anthropic and Hugging Face CEOs Call for Global AI Regulation at UN Security Council.
- On September 24–25, the White House reportedly asked OpenAI and Anthropic to delay sharing new models with UK testers until the models have gone through U.S. government review first — see White House Asks OpenAI & Anthropic to Delay Sharing Models with UK Testers Until U.S. Review.
The ExploitBench Critical rating sits inside that cluster: a frontier lab with a model that can do serious cyber work, a UN forum asking for international coordination, and a White House asserting a U.S.-first review pipeline. The model's capabilities are not waiting for the governance to catch up.
What we do not know yet
- Whether "Critical" on ExploitBench is a formal OpenAI-internal classification with written criteria, or a looser editorial label on a report.
- Whether OpenAI has changed any deployment controls, rate limits, or access rules for GPT-6 Astra in response to the rating, or whether the rating is observational only.
- Whether the rating triggers any obligation under the U.S. voluntary frameworks OpenAI has agreed to, or under the UK's incoming AI safety regime, which has been building toward its own test-and-evaluate pipeline.
- Whether other frontier models — Anthropic's Claude family, Google's Gemini — have been rated on the same ExploitBench scale, and where they stand. OpenAI's test is its own; cross-lab comparison requires a shared benchmark.
Why this is worth watching
The Sep 4 score told us GPT-6 Astra can do the work. The Sep 25 rating tells us the people who built the test now think that work is critical-risk.
That sequencing — score first, classification later — is likely to become a recurring pattern as frontier models improve on the cyber evaluations that labs use to measure them. A model clears a threshold on Tuesday; the people measuring it decide what that threshold means on Friday; and in the gap between those two days, the model has been in the wild the whole time.
Sources: tech-insider.org (September 25, 2026); The Hacker News (September 4, 2026); Yellow.com (September 4, 2026); Gagadget.com (September 4, 2026); CryptoRank (September 4, 2026); OpenAI "Path to Astra: critical capabilities and frontier safeguards" (September 1, 2026).