OpenAI launched GPT-6 Astra on September 3, 2026, and called it the most intelligent and aligned model in the world. The GPT-6 Astra benchmarks behind that claim are genuinely impressive in places and genuinely oversold in others. If you are trying to decide whether Astra is the model for your next project, the scores only help once you know who produced them and under what settings.
This guide separates the vendor numbers from the independent ones, explains what each benchmark measures, and tells you what a jump or drop actually means when you sit down to build. Emergent covered the release in its launch report, and here we go deeper on the numbers.
How to read GPT-6 Astra's benchmark numbers
Benchmark scores mean nothing until you know who ran the test and how. A model's published results fall into three buckets, and mixing them is how people end up believing a launch is bigger than it is.
The three tiers you will see across every GPT-6 Astra benchmark discussion:
- Vendor-reported: OpenAI ran the evaluation and chose the settings. These numbers are directional, not neutral. OpenAI notes its scores reflect the maximum at any effort level, which flatters results but raises latency and token use.
- Third-party verified: an independent lab such as Artificial Analysis runs the same models through one harness. These are the numbers where cross-model comparison actually holds.
- External vendor: a rival lab reports its own model's score, as Meta did for Muse Spark 1.3. Useful context, but run on a different setup.
Comparison is only safe when the benchmark, the harness, and the effort setting match. When they do not, a table can show two scores side by side that were never measured the same way. OpenAI's own footnotes admit several of these mismatches, and we flag them below rather than pretending the rows line up.
What GPT-6 Astra scores on the headline benchmarks
GPT-6 Astra posts the highest scores OpenAI has ever published on math, cybersecurity, and computer use, all listed in its official launch tables. These are the results driving the launch coverage, and every figure here is vendor-reported, so read them as OpenAI's best case rather than a neutral ranking.
1. FrontierMath Tier 4: 97.6%
On FrontierMath Tier 4, a set of research-level math problems built to resist AI, Astra scored 97.6%. Saturating a test designed to stay ahead of models signals real reasoning gains, though frontier math rarely maps to the everyday tasks most people run.
2. ExploitBench: 100%
On ExploitBench, which measures whether a model can turn a known software vulnerability into a working exploit, Astra hit a perfect 100%, up from Sol's 78.5%. That jump is why OpenAI says Astra is the first model to cross the Critical threshold under its own Preparedness Framework, a vendor safety classification rather than an industry-standard rating, and why advanced cyber features ship gated behind its Daybreak program. Emergent's cyber-threshold report covers what that gating means in practice.
3. OSWorld 2.0: 72.6% at half the time
On OSWorld 2.0, which tests whether an agent can complete real desktop tasks like navigating apps and moving files, Astra scored 72.6% while taking roughly 47% less time per task than Sol. For anyone delegating actual computer work, the time drop matters more than the score. In OpenAI's latency simulation on the OSWorld 2.0 offline subset, Astra finished the average task in about 40 minutes versus Sol's 75.
Table 1 - Selected vendor-reported GPT-6 Astra benchmark highlights
The ARC-AGI-3 99.9% comes with an asterisk
Astra's headline 99.9% on ARC-AGI-3 is real, but it depends on a custom harness most users will never touch. ARC-AGI-3 tests whether a model can learn to solve interactive environments it has never seen, which is why near-saturation grabbed so much attention.
The catch sits in OpenAI's own footnote. Astra ran on a modified Responses API harness that changed two settings to better match real-world performance. Secondary coverage from outlets including DataCamp reported that stateless API calls, the kind a normal integration makes, score lower than the harness result. In a quote hosted on OpenAI's launch page, ARC Prize Foundation's Greg Kamradt described Astra as surpassing the human action-efficiency baseline on 96% of levels, effectively reaching human parity, rather than solving general intelligence.
What this means for you: you will not see 99.9% behavior from a plain API call. The number reflects a tuned setup, so treat it as a ceiling under ideal conditions rather than the performance you should expect in production.
Where independent testing tells a different story
Independent testing puts GPT-6 Astra in a far more modest position than the launch tables suggest. Artificial Analysis, which runs every major model through one neutral harness, found Astra roughly level with its own predecessor on general intelligence.
1. Intelligence Index: tied with its own predecessor
On the Artificial Analysis Intelligence Index, which aggregates reasoning, knowledge, and coding evaluations into one score, Astra landed at 61.2. That ties GPT-5.6 Sol at 60.9 and trails Claude Fable 5.1 at 65.7. Tying your own previous model on the neutral aggregate is not the profile of a generational leap, and it is the single most important counterweight to the AGI framing.
2. Coding Agent Index: the win is cost, not capability
The coding picture is better, but the story is cost, not raw capability. On the Artificial Analysis Coding Agent Index, Astra scored 67.0, roughly level with the Fable models, while matching Fable 5's score at less than half the cost per task thanks to large token-efficiency gains. On the coding harness at maximum effort, Astra used about a third of the tokens Sol needed. Anthropic's own tiers split on a similar cost-versus-capability line, which our Sonnet vs Opus comparison breaks down.
3. Hallucination: halved, but still high
Astra also cut hallucination sharply. On Artificial Analysis's knowledge test, its hallucination rate fell from 92% to 51% at maximum effort, without sacrificing accuracy. That is a real improvement, yet a 51% rate is still high, so you should not trust unverified factual recall from Astra any more than you would from a careful but fallible assistant.
On the neutral Intelligence Index, GPT-6 Astra scores 61.2, statistically tied with the model it replaces and five points behind Claude Fable 5.1.
How GPT-6 Astra's scores compare
Placed against its closest rivals, GPT-6 Astra wins on math and cyber, ties on the neutral coding index, and trails Fable 5.1 on the Intelligence Index and Humanity's Last Exam. The table below is a reference for reading Astra's numbers in context, not a verdict on which model to buy.
If you are weighing model families more broadly, our ChatGPT vs Claude vs Gemini guide covers the wider field. Rows are labeled by source tier, and the coding and intelligence rows drawn from Artificial Analysis are the ones where comparison is safest.
A few rows deserve a caution. OpenAI's OSWorld figure for Claude uses a different release than Anthropic's own system card, so those cells are not strictly comparable. Several Claude columns on the science and cyber tests are blank because the models refused the evaluations, which OpenAI notes in its footnotes. We have kept only rows where the comparison is reasonably clean.
Table 2 - GPT-6 Astra benchmark comparison against Sol, Fable 5.1, and Opus 5
Two rows carry the honest caveats. On DeepSWE v1.1, an agentic coding test, Astra's 74.1% sits within a point of Opus 5 at 73.7%, close enough to call a tie, and Meta reported a higher 75.4% for Muse Spark 1.3 at maximum reasoning. Even Gemini 3.8 Flash, a smaller Flash-class model, lands at 73.8% on the same test, so Astra's agentic-coding lead is thin. On Humanity's Last Exam with tools, Astra's 57.2% clearly trails Fable 5.1's 65.0%. For the full head-to-head on the Anthropic side, see the Fable 5.1 versus Opus 5 breakdown.
GPT-6 Astra pricing and how to access it
GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, which is 2.5 times the price of GPT-5.6 Sol and roughly on par with Claude Fable 5.1. Cache reads carry a 90% discount and cache writes a 25% premium. A Fast mode offers up to twice the speed at twice the standard price.
The sticker price overstates the real cost gap on coding, but not everywhere. Because Astra uses far fewer tokens to finish a task, Artificial Analysis found its cost per completed coding task can fall below rivals even at the higher rate. On general work the opposite holds: Artificial Analysis found Astra about 75% more expensive per Intelligence Index task than Sol at maximum effort, since the 10% token saving there does not offset the 2.5 times price increase. What you pay per token and what you pay per finished job are not the same number.
Table 3 - GPT-6 Astra pricing as of September 2026
On availability, Astra rolled out first to a limited set of organizations, then to all ChatGPT Plus, Pro, Business, and Enterprise users over the following days. Developers reach it through the OpenAI API as gpt-6-astra, plus Microsoft Azure and AWS Bedrock. Enterprise access is off by default at launch, and Pro, Business, and Enterprise plans also unlock GPT-6 Astra Pro.
What GPT-6 Astra's benchmarks actually tell you
GPT-6 Astra is a genuine leap in efficiency and computer use over OpenAI's own line, not a decisive jump past the best models from rivals. Read the benchmarks that way and the launch makes sense: Astra separates itself on the messy, agentic, tool-using tasks and on coding cost per task, while sitting level with the field on raw neutral intelligence.
The gains that hold up under scrutiny are the practical ones. Astra does more work in less time on real computer tasks, spends far fewer tokens to reach a given result, and hallucinates about half as often as Sol. Those are the numbers a builder feels day to day, more than a saturated math score.
Where rivals still win is worth stating plainly. Claude Fable 5.1 leads the neutral Intelligence Index and takes Humanity's Last Exam. On agentic coding, Opus 5 and Meta's Muse Spark 1.3 are within noise of Astra. The right benchmark to care about is the one that matches your job, so a coding-agent score should weigh more than an abstract-reasoning score if you are shipping software.
On the AGI question
The AGI talk around Astra rests almost entirely on the saturated benchmarks, and those are the ones with the biggest asterisks. OpenAI called Astra the most intelligent model in the world, and Greg Brockman told reporters it is not unreasonable to feel we are now in the AGI era. That framing leans on ARC-AGI-3 and FrontierMath, both harness-sensitive or narrow.
The benchmark built specifically to resist saturation, the Artificial Analysis Intelligence Index, is exactly where Astra looks incremental and lands behind Fable 5.1. There is no agreed definition of AGI, and no single benchmark settles it. Astra is a strong model that tops specific tests under specific conditions, which is a claim the evidence supports. Whether that amounts to the AGI era is a question of interpretation, not a settled result.
Build your next app on Emergent
The honest takeaway on GPT-6 Astra benchmarks is that setup and source decide the story. Astra is state-of-the-art on math, cyber, and computer use in OpenAI's own tables, incremental on the neutral aggregate, and genuinely ahead on coding efficiency and cost per coding task. It is a serious model with real strengths and a few clear losses to Claude Fable 5.1, and the AGI framing runs ahead of what the neutral benchmarks show. Which model wins comes down to the job in front of you, so the ideal setup lets you switch as the task changes.
That is exactly how building on Emergent works. You can use GPT models, Claude models, or other frontier-lab models, and every project runs through a single Universal LLM Key, so you choose the model that fits the job without wiring up separate accounts. Describe the software you want, pick your model, and let Emergent build the full-stack app that runs your business.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







