HomeLearn

GPT-6 Astra Benchmarks: What the Numbers Actually Show

GPT-6 Astra benchmarks show state-of-the-art math and cyber scores, but independent tests place it behind Claude Fable 5.1. See what the numbers mean.

Bhavyadeep Sinh Rathod
Written by
Bhavyadeep
Priyanka Singh
Reviewed by
Priyanka Singh
Published: 
Sep 4, 2026
0
 min read
Table of Contents

TL;DR

  • GPT-6 Astra benchmarks are strongest where OpenAI ran the tests: it saturates FrontierMath Tier 4, ARC-AGI-3, and ExploitBench in OpenAI's own tables.
  • On the independent Artificial Analysis Intelligence Index, Astra scores 61.2, effectively tied with its predecessor GPT-5.6 Sol and behind Claude Fable 5.1 at 65.7.
  • Astra ties Fable 5.1 on the neutral coding index and trails only in Artificial Analysis's Claude Code setup, while losing Humanity's Last Exam to Fable 5.1, so the launch is not a clean sweep.
  • The real gain is coding efficiency: Astra matches Fable 5's coding-agent score at less than half the cost per task and cuts hallucination roughly in half.
  • Pricing is $10 per million input tokens and $50 per million output, which is 2.5 times Sol and on par with Fable 5.1.

OpenAI launched GPT-6 Astra on September 3, 2026, and called it the most intelligent and aligned model in the world. The GPT-6 Astra benchmarks behind that claim are genuinely impressive in places and genuinely oversold in others. If you are trying to decide whether Astra is the model for your next project, the scores only help once you know who produced them and under what settings.

This guide separates the vendor numbers from the independent ones, explains what each benchmark measures, and tells you what a jump or drop actually means when you sit down to build. Emergent covered the release in its launch report, and here we go deeper on the numbers.

How to read GPT-6 Astra's benchmark numbers

Benchmark scores mean nothing until you know who ran the test and how. A model's published results fall into three buckets, and mixing them is how people end up believing a launch is bigger than it is.

The three tiers you will see across every GPT-6 Astra benchmark discussion:

  • Vendor-reported: OpenAI ran the evaluation and chose the settings. These numbers are directional, not neutral. OpenAI notes its scores reflect the maximum at any effort level, which flatters results but raises latency and token use.
  • Third-party verified: an independent lab such as Artificial Analysis runs the same models through one harness. These are the numbers where cross-model comparison actually holds.
  • External vendor: a rival lab reports its own model's score, as Meta did for Muse Spark 1.3. Useful context, but run on a different setup.

Comparison is only safe when the benchmark, the harness, and the effort setting match. When they do not, a table can show two scores side by side that were never measured the same way. OpenAI's own footnotes admit several of these mismatches, and we flag them below rather than pretending the rows line up.

What GPT-6 Astra scores on the headline benchmarks

GPT-6 Astra posts the highest scores OpenAI has ever published on math, cybersecurity, and computer use, all listed in its official launch tables. These are the results driving the launch coverage, and every figure here is vendor-reported, so read them as OpenAI's best case rather than a neutral ranking.

1. FrontierMath Tier 4: 97.6%

On FrontierMath Tier 4, a set of research-level math problems built to resist AI, Astra scored 97.6%. Saturating a test designed to stay ahead of models signals real reasoning gains, though frontier math rarely maps to the everyday tasks most people run.

2. ExploitBench: 100%

On ExploitBench, which measures whether a model can turn a known software vulnerability into a working exploit, Astra hit a perfect 100%, up from Sol's 78.5%. That jump is why OpenAI says Astra is the first model to cross the Critical threshold under its own Preparedness Framework, a vendor safety classification rather than an industry-standard rating, and why advanced cyber features ship gated behind its Daybreak program. Emergent's cyber-threshold report covers what that gating means in practice.

3. OSWorld 2.0: 72.6% at half the time

On OSWorld 2.0, which tests whether an agent can complete real desktop tasks like navigating apps and moving files, Astra scored 72.6% while taking roughly 47% less time per task than Sol. For anyone delegating actual computer work, the time drop matters more than the score. In OpenAI's latency simulation on the OSWorld 2.0 offline subset, Astra finished the average task in about 40 minutes versus Sol's 75.

Table 1 - Selected vendor-reported GPT-6 Astra benchmark highlights

Benchmark What it tests GPT-6 Astra
FrontierMath Tier 4 (v2) Research-level math 97.6%
ExploitBench Building working exploits 100.0%
OSWorld 2.0 (offline) Real desktop tasks 72.6%
Agents' Last Exam Complex professional work 59.3%
SRE-Bench (88.0% in 1 try, 99.2% in 4) Reverse-engineering binaries 99.2%

The ARC-AGI-3 99.9% comes with an asterisk

Astra's headline 99.9% on ARC-AGI-3 is real, but it depends on a custom harness most users will never touch. ARC-AGI-3 tests whether a model can learn to solve interactive environments it has never seen, which is why near-saturation grabbed so much attention.

The catch sits in OpenAI's own footnote. Astra ran on a modified Responses API harness that changed two settings to better match real-world performance. Secondary coverage from outlets including DataCamp reported that stateless API calls, the kind a normal integration makes, score lower than the harness result. In a quote hosted on OpenAI's launch page, ARC Prize Foundation's Greg Kamradt described Astra as surpassing the human action-efficiency baseline on 96% of levels, effectively reaching human parity, rather than solving general intelligence.

What this means for you: you will not see 99.9% behavior from a plain API call. The number reflects a tuned setup, so treat it as a ceiling under ideal conditions rather than the performance you should expect in production.

Where independent testing tells a different story

Independent testing puts GPT-6 Astra in a far more modest position than the launch tables suggest. Artificial Analysis, which runs every major model through one neutral harness, found Astra roughly level with its own predecessor on general intelligence.

1. Intelligence Index: tied with its own predecessor

On the Artificial Analysis Intelligence Index, which aggregates reasoning, knowledge, and coding evaluations into one score, Astra landed at 61.2. That ties GPT-5.6 Sol at 60.9 and trails Claude Fable 5.1 at 65.7. Tying your own previous model on the neutral aggregate is not the profile of a generational leap, and it is the single most important counterweight to the AGI framing.

2. Coding Agent Index: the win is cost, not capability

The coding picture is better, but the story is cost, not raw capability. On the Artificial Analysis Coding Agent Index, Astra scored 67.0, roughly level with the Fable models, while matching Fable 5's score at less than half the cost per task thanks to large token-efficiency gains. On the coding harness at maximum effort, Astra used about a third of the tokens Sol needed. Anthropic's own tiers split on a similar cost-versus-capability line, which our Sonnet vs Opus comparison breaks down.

3. Hallucination: halved, but still high

Astra also cut hallucination sharply. On Artificial Analysis's knowledge test, its hallucination rate fell from 92% to 51% at maximum effort, without sacrificing accuracy. That is a real improvement, yet a 51% rate is still high, so you should not trust unverified factual recall from Astra any more than you would from a careful but fallible assistant.

On the neutral Intelligence Index, GPT-6 Astra scores 61.2, statistically tied with the model it replaces and five points behind Claude Fable 5.1.

How GPT-6 Astra's scores compare

Placed against its closest rivals, GPT-6 Astra wins on math and cyber, ties on the neutral coding index, and trails Fable 5.1 on the Intelligence Index and Humanity's Last Exam. The table below is a reference for reading Astra's numbers in context, not a verdict on which model to buy.

If you are weighing model families more broadly, our ChatGPT vs Claude vs Gemini guide covers the wider field. Rows are labeled by source tier, and the coding and intelligence rows drawn from Artificial Analysis are the ones where comparison is safest.

A few rows deserve a caution. OpenAI's OSWorld figure for Claude uses a different release than Anthropic's own system card, so those cells are not strictly comparable. Several Claude columns on the science and cyber tests are blank because the models refused the evaluations, which OpenAI notes in its footnotes. We have kept only rows where the comparison is reasonably clean.

Table 2 - GPT-6 Astra benchmark comparison against Sol, Fable 5.1, and Opus 5

Benchmark (source) Astra GPT-5.6 Sol Fable 5.1 Opus 5
Intelligence Index (Artificial Analysis) 61.2 60.9 65.7 63.1
FrontierMath Tier 4 (OpenAI) 97.6% 83.0% 87.8% 73.2%
Terminal-Bench 4.0 (OpenAI) 57.9% 37.3% 55.8% 52.3%
DeepSWE v1.1 (OpenAI) 74.1% 72.7% 67.4% 73.7%
Humanity's Last Exam, tools (OpenAI) 57.2% n/a 65.0% 63.6%

Two rows carry the honest caveats. On DeepSWE v1.1, an agentic coding test, Astra's 74.1% sits within a point of Opus 5 at 73.7%, close enough to call a tie, and Meta reported a higher 75.4% for Muse Spark 1.3 at maximum reasoning. Even Gemini 3.8 Flash, a smaller Flash-class model, lands at 73.8% on the same test, so Astra's agentic-coding lead is thin. On Humanity's Last Exam with tools, Astra's 57.2% clearly trails Fable 5.1's 65.0%. For the full head-to-head on the Anthropic side, see the Fable 5.1 versus Opus 5 breakdown.

GPT-6 Astra pricing and how to access it

GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, which is 2.5 times the price of GPT-5.6 Sol and roughly on par with Claude Fable 5.1. Cache reads carry a 90% discount and cache writes a 25% premium. A Fast mode offers up to twice the speed at twice the standard price.

The sticker price overstates the real cost gap on coding, but not everywhere. Because Astra uses far fewer tokens to finish a task, Artificial Analysis found its cost per completed coding task can fall below rivals even at the higher rate. On general work the opposite holds: Artificial Analysis found Astra about 75% more expensive per Intelligence Index task than Sol at maximum effort, since the 10% token saving there does not offset the 2.5 times price increase. What you pay per token and what you pay per finished job are not the same number.

Table 3 - GPT-6 Astra pricing as of September 2026

Plan detail GPT-6 Astra
Input (per million tokens) $10
Output (per million tokens) $50
Cache read discount 90%
Fast mode 2x speed at 2x price

On availability, Astra rolled out first to a limited set of organizations, then to all ChatGPT Plus, Pro, Business, and Enterprise users over the following days. Developers reach it through the OpenAI API as gpt-6-astra, plus Microsoft Azure and AWS Bedrock. Enterprise access is off by default at launch, and Pro, Business, and Enterprise plans also unlock GPT-6 Astra Pro.

What GPT-6 Astra's benchmarks actually tell you

GPT-6 Astra is a genuine leap in efficiency and computer use over OpenAI's own line, not a decisive jump past the best models from rivals. Read the benchmarks that way and the launch makes sense: Astra separates itself on the messy, agentic, tool-using tasks and on coding cost per task, while sitting level with the field on raw neutral intelligence.

The gains that hold up under scrutiny are the practical ones. Astra does more work in less time on real computer tasks, spends far fewer tokens to reach a given result, and hallucinates about half as often as Sol. Those are the numbers a builder feels day to day, more than a saturated math score.

Where rivals still win is worth stating plainly. Claude Fable 5.1 leads the neutral Intelligence Index and takes Humanity's Last Exam. On agentic coding, Opus 5 and Meta's Muse Spark 1.3 are within noise of Astra. The right benchmark to care about is the one that matches your job, so a coding-agent score should weigh more than an abstract-reasoning score if you are shipping software.

On the AGI question

The AGI talk around Astra rests almost entirely on the saturated benchmarks, and those are the ones with the biggest asterisks. OpenAI called Astra the most intelligent model in the world, and Greg Brockman told reporters it is not unreasonable to feel we are now in the AGI era. That framing leans on ARC-AGI-3 and FrontierMath, both harness-sensitive or narrow.

The benchmark built specifically to resist saturation, the Artificial Analysis Intelligence Index, is exactly where Astra looks incremental and lands behind Fable 5.1. There is no agreed definition of AGI, and no single benchmark settles it. Astra is a strong model that tops specific tests under specific conditions, which is a claim the evidence supports. Whether that amounts to the AGI era is a question of interpretation, not a settled result.

Build your next app on Emergent

The honest takeaway on GPT-6 Astra benchmarks is that setup and source decide the story. Astra is state-of-the-art on math, cyber, and computer use in OpenAI's own tables, incremental on the neutral aggregate, and genuinely ahead on coding efficiency and cost per coding task. It is a serious model with real strengths and a few clear losses to Claude Fable 5.1, and the AGI framing runs ahead of what the neutral benchmarks show. Which model wins comes down to the job in front of you, so the ideal setup lets you switch as the task changes.

That is exactly how building on Emergent works. You can use GPT models, Claude models, or other frontier-lab models, and every project runs through a single Universal LLM Key, so you choose the model that fits the job without wiring up separate accounts. Describe the software you want, pick your model, and let Emergent build the full-stack app that runs your business.

Start Building on Emergent.

Was this article helpful?
About the writer
Bhavyadeep
Bhavyadeep Sinh Rathod
Content Manager

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free

Frequently Asked Questions

Your Questions, Answered

Is GPT-6 Astra better than Claude Fable 5.1?
It depends on the task. GPT-6 Astra leads on math, cybersecurity, and computer-use benchmarks in OpenAI's tables. Claude Fable 5.1 leads the independent Artificial Analysis Intelligence Index at 65.7 versus Astra's 61.2, and beats Astra on Humanity's Last Exam with tools. On those two measures Fable 5.1 has the edge, while Astra wins on agentic coding efficiency and cost per coding task.
What is ARC-AGI-3?
ARC-AGI-3 is a benchmark that tests whether an AI model can learn to solve interactive environments it has never encountered before. It is designed to resist saturation, so a high score signals a model can navigate novel problems rather than recall training data. GPT-6 Astra scored 99.9%, though that result used a custom harness rather than a standard API call.
Why is GPT-6 Astra's ARC-AGI-3 score disputed?
The 99.9% depends on a tuned setup. OpenAI ran Astra on a modified Responses API harness that changed two settings, and secondary coverage reported that stateless API calls, the kind normal integrations use, score lower. The number is real but reflects best-case conditions, so it does not represent the performance you would see from a standard production call.
How much does GPT-6 Astra cost?
GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens as of September 2026. That is 2.5 times the price of GPT-5.6 Sol and roughly on par with Claude Fable 5.1. Cache reads get a 90% discount. Because Astra uses fewer tokens per task, its cost per completed job can still land below rivals despite the higher rate.
Is GPT-6 Astra AGI?
No single benchmark settles that, and there is no agreed definition of AGI. GPT-6 Astra saturates specific tests like FrontierMath and ARC-AGI-3, which fuels the AGI talk, but those results are narrow or harness-sensitive. On the neutral Artificial Analysis Intelligence Index, Astra is incremental and sits behind Claude Fable 5.1. It is a strong model, but the AGI label is interpretation, not a settled result.
Is GPT-6 Astra available on the API?
Yes. GPT-6 Astra is available through the OpenAI API as gpt-6-astra, as well as Microsoft Azure and AWS Bedrock. It rolled out first to a limited set of organizations, then to all ChatGPT Plus, Pro, Business, and Enterprise users. Enterprise access is off by default at launch, and higher tiers also unlock GPT-6 Astra Pro.
Start Building
on Emergent today
Try Emergent
This is some text inside of a div block.
This is some text inside of a div block.
Note

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

https://api.linear.app/graphql