HomeLearn

GPT-6.1 Sol Benchmarks: Near-Astra Scores, Lower Cost

GPT-6.1 Sol benchmarks: 75.2% on DeepSWE, 52 on the Artificial Analysis index, and $0.72 per task. See where it matches GPT-6 Astra and where it falls short.

Bhavyadeep
Written by
Bhavyadeep
Sakthy
Reviewed by
Sakthy
Last updated: 
September 30, 2026
0
 min read
Select Emergent as your Preferred news source
Table of Contents

TL;DR

  • GPT-6.1 Sol benchmarks land close to the flagship: it scores 52 on the Artificial Analysis Intelligence Index, 1 point behind GPT-6 Astra, at $0.72 per task versus Astra's $3.26.
  • On OpenAI's DeepSWE v1.1 coding test, it reaches 75.2% at high effort, above Astra's best run of 74.1%, for $0.65 per task instead of $4.43.
  • The biggest jump over GPT-6 Sol comes at low effort, where DeepSWE rises from 37.2% to 64.4% for about the same cost.
  • Astra still leads on hard science tasks, and Claude Opus 5.5 and Sonnet 5.5 both score higher on the Artificial Analysis index.
  • Verdict: 6.1 Sol is the new default Sol for agents and automations, and high effort, not max, is its sweet spot.

‍

GPT-6.1 Sol is a point release that changes which model you should pick. OpenAI shipped it at DevDay on September 29, 2026, seven days after GPT-6 Sol, and the GPT-6.1 Sol benchmarks put it within a point or two of GPT-6 Astra for a fraction of the cost.

If you are choosing a model for an agent or an app that runs daily, near-flagship quality at mid-tier prices changes the math.

The catch is that most headline numbers come from OpenAI itself. Artificial Analysis has since confirmed the core claim independently, and it also found a trade-off OpenAI did not mention. For the baseline this release supersedes, see our GPT-6 Sol benchmarks breakdown.

GPT-6.1 Sol benchmarks at a glance

GPT-6.1 Sol ties or beats Astra on coding and trails it by one to two points almost everywhere else, while costing far less per task. The table separates vendor-reported results from independent ones, because the two are not measured the same way.

Table 1 - GPT-6.1 Sol benchmark summary by source, as of September 2026.

Benchmark GPT-6.1 Sol Comparison Source
Intelligence Index (max effort) 52 GPT-6 Sol: 48. GPT-6 Astra: 1 point higher Artificial Analysis (independent)
Cost per Intelligence Index task (max) $0.72 GPT-6 Astra: $3.26. GPT-6 Sol: $1.05 Artificial Analysis (independent)
Coding Agent Index (xhigh) 1 point above Astra At less than 15% of Astra's cost per task Artificial Analysis (independent)
DeepSWE v1.1 (high effort) 75.2% GPT-6 Astra best run: 74.1% OpenAI (vendor-reported, chart data)
OSWorld 2.0 offline (max) Within 2.1 points of Astra 7 points above GPT-6 Sol OpenAI (vendor-reported)
AutomationBench 1.0.6 (medium) 2.2 points above Opus 5.5 4.8 points above GPT-6 Sol OpenAI (vendor-reported)
Terminal-Bench Science 0.1 (max) More than double GPT-6 Sol GPT-6 Astra leads at 68.1% OpenAI (vendor-reported)
Factual error rate (low effort) 7.7% GPT-6 Sol: 11.4% OpenAI (vendor-reported)
API price per 1M tokens (input, cache read, output) $2, $0.10, $10 GPT-6 Astra: $10, $1, $50 OpenAI

Weight the independent rows most. Artificial Analysis runs every model on the same harness, while OpenAI's rows show where the company chose to test.

What changed from GPT-6 Sol

GPT-6.1 Sol is a straight quality upgrade at the same list price, with cheaper cache reads and one real trade-off in speed. OpenAI positions it as an upgrade to GPT-6 Sol, launched a week after that model, which our GPT-6.1 Sol launch coverage tracks.

The gains show up in nearly every category:

  • Overall intelligence: Artificial Analysis measured a 4-point gain on its Intelligence Index, from 48 to 52.
  • Coding: DeepSWE v1.1 improved by 6.4 points at its best setting, and Terminal-Bench 4.0 jumped 12 points in independent testing.
  • Knowledge work: GDP.pdf rose 6 points and AA-Briefcase gained about 80 Elo, per Artificial Analysis.
  • Accuracy: Hallucination rate on AA-Omniscience fell from 60% to 54%.

Pricing stayed at $2 per 1M input tokens and $10 per 1M output tokens. Cache reads dropped from $0.20 to $0.10 per 1M tokens. That is a 50% cut from GPT-6 Sol and a 95% discount on 6.1 Sol's own standard input rate. That cut matters most for agents that reread the same long context on every step. The full rate card, including long-context tiers, is in our GPT-6 Sol pricing guide.

Two changes cut the other way. GPT-6.1 Sol uses 10% to 30% more output tokens than GPT-6 Sol, and Artificial Analysis clocks it at 67.8 output tokens per second against 104 for GPT-6 Sol. It also drops the "none" reasoning setting. OpenAI's migration guide says to use "low" instead, so apps that set "none" explicitly need a config change and a quick regression test.

How GPT-6.1 Sol scores on OpenAI's evaluations

OpenAI's results show the largest gains in coding and computer use, and every result is framed against cost per task rather than score alone. Most results in the OpenAI announcement appear as charts with relative claims. The DeepSWE figures below come from that chart's underlying data.

1. Coding: DeepSWE v1.1

GPT-6.1 Sol scores 75.2% on DeepSWE v1.1 at high effort, the highest point on OpenAI's DeepSWE chart across all three models and five effort settings. DeepSWE measures long, original software engineering tasks in real codebases, so it is a fair proxy for coding agents.

Table 2 - DeepSWE v1.1 score and cost per task by reasoning effort (OpenAI, vendor-reported, as of September 2026).

Effort GPT-6.1 Sol GPT-6 Sol GPT-6 Astra
Low 64.4% at $0.17 37.2% at $0.16 67.0% at $1.60
Medium 73.0% at $0.42 56.6% at $0.38 72.8% at $3.08
High 75.2% at $0.65 65.3% at $0.64 73.2% at $3.92
Xhigh 71.9% at $0.79 66.6% at $1.00 74.1% at $4.43
Max 71.9% at $1.57 68.8% at $2.74 73.2% at $7.50

At medium effort, 6.1 Sol edges out Astra's medium run for about a seventh of the cost. Its score also drops after high effort, so xhigh and max buy nothing here.

2. Professional work: GDP.pdf and AutomationBench

GPT-6.1 Sol beats Claude Opus 5.5 on both of OpenAI's professional work tests. On GDP.pdf, which asks models to answer questions from complex PDFs in finance, healthcare, legal, and other fields, it outscores Opus 5.5 at less than half the cost per task. It also comes close to Astra's leading score at about a fifth of the cost.

On AutomationBench 1.0.6, agents complete end-to-end business workflows using 47 tools across sales, operations, support, and HR. GPT-6.1 Sol scores 2.2 points above Opus 5.5 at medium effort, at about a third of the cost. That is also 4.8 points better than GPT-6 Sol at the same setting.

3. Computer use: OSWorld 2.0

GPT-6.1 Sol gains 7 points over GPT-6 Sol on OSWorld 2.0's offline set at max effort, at less than half the cost. It lands within 2.1 points of Astra at about one-seventh of Astra's cost per task. OSWorld tests long, multi-app workflows, the kind of work browser and desktop agents do.

4. Scientific research: Terminal-Bench Science 0.1

Science is where GPT-6.1 Sol improves most and where Astra keeps its clearest lead. On Terminal-Bench Science 0.1, which covers data analysis, simulation, and model fitting, 6.1 Sol more than doubles GPT-6 Sol's score at max effort.

It costs $5.47 per task at max effort, compared with $23.21 for Opus 5.5 and $23.80 for Astra. Astra still posts the top score at 68.1%, and OpenAI itself recommends Astra for the most difficult scientific research tasks.

5. Factuality and safety

GPT-6.1 Sol makes fewer factual errors, especially at low effort. On deliberately difficult prompts drawn from conversations where users flagged a past error, the share of answers with a mistake fell from 11.4% to 7.7%. Across settings, its error rate stayed within 1.9 points of Astra's.

Safety results moved the same way. When an agent's search tool is broken, GPT-6.1 Sol fails to tell the user in 2.1% of cases, versus 4.9% for GPT-6 Sol and 1.5% for Astra.

What Artificial Analysis measured independently

Independent testing confirms the near-Astra claim. In its launch analysis, Artificial Analysis scores GPT-6.1 Sol at 52 on its Intelligence Index, 1 point below GPT-6 Astra. It also found no cheaper model at that level of intelligence when it ran the tests.

The cost gap is the headline. At max effort, a task on the Intelligence Index costs $0.72 with 6.1 Sol, $1.05 with GPT-6 Sol, and $3.26 with Astra. Every effort level of 6.1 Sol sat on Artificial Analysis's intelligence-versus-cost frontier at the time of testing.

[SCREENSHOT: Artificial Analysis Intelligence Index vs. Cost per Task chart, showing GPT-6.1 Sol effort levels against GPT-6 Astra and GPT-6 Sol]

Artificial Analysis Intelligence Index vs. cost per task for GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Astra Alt text - GPT-6.1 Sol benchmarks chart comparing Intelligence Index score and cost per task with GPT-6 Astra

Coding tells a similar story. On the Coding Agent Index, 6.1 Sol at max effort gains 3 points on GPT-6 Sol and sits 2 points below Astra. At xhigh effort it scores 1 point above Astra for less than 15% of the cost, and it outperforms its own max setting by 3 points.

Speed is the weak spot. The model page lists 67.8 output tokens per second at max effort and rates it slower than average among the models it tracks. Throughput is not end-to-end latency, so test in chat-style apps before switching.

Which reasoning effort to use

Start at high effort for coding agents and medium for business workflows, and skip max unless your own tests prove it helps. Both OpenAI's DeepSWE data and Artificial Analysis's coding results show 6.1 Sol peaking below its highest setting.

GPT-6.1 Sol supports low, medium, high, xhigh, and max effort, with medium as the default, according to OpenAI's model guide. A practical starting point:

  • Low: Scores 64.4% on DeepSWE at $0.17 per task, close to Astra's low setting at a tenth of the cost. Use it for high-volume tasks like tagging, routing, or short summaries.
  • Medium: This is where OpenAI reports its AutomationBench win over Opus 5.5. Use it for multi-step business workflows.
  • High: Posts the highest score on OpenAI's DeepSWE chart. Use it for coding agents and complex builds.
  • Xhigh or max: Worth testing only for hard research or long computer-use runs, where OpenAI's max-effort results show the largest gains.

GPT-6.1 Sol vs GPT-6 Astra: where the gap remains

Astra is still the model for the hardest scientific research tasks, and elsewhere the gap is small enough that most builders will not notice it. Astra keeps a clear lead only on Terminal-Bench Science 0.1, where it scores 68.1%.

Everywhere else the margin is narrow. 6.1 Sol trails by 1 point on the Intelligence Index, 2 points on the Coding Agent Index at max effort, and 2.1 points on OSWorld 2.0. On DeepSWE it wins outright.

The price gap is not narrow. Astra lists at $10 input and $50 output per 1M tokens, five times Sol's rates, and costs about 4.5 times as much per Intelligence Index task. Our Sol and Luna explainer covers how the three GPT-6 tiers fit together.

How GPT-6.1 Sol compares with Claude Opus 5.5 and Sonnet 5.5

The answer depends on who is measuring. OpenAI's chosen tests show GPT-6.1 Sol beating Claude Opus 5.5, while the broader independent index puts both Claude models ahead.

Table 3 - GPT-6.1 Sol vs Claude models by source, as of September 2026.

Measure GPT-6.1 Sol Claude result Source
AutomationBench 1.0.6 (medium) 2.2 points higher Opus 5.5, at about 3x the cost OpenAI (vendor-reported)
GDP.pdf Higher score Opus 5.5, at over 2x the cost OpenAI (vendor-reported)
Terminal-Bench Science cost (max) $5.47 per task Opus 5.5: $23.21 per task OpenAI (vendor-reported)
Intelligence Index 52 Sonnet 5.5: 56. Opus 5.5 (max): 58 Artificial Analysis (independent)

Artificial Analysis scores Claude Sonnet 5.5 at 56, 4 points above 6.1 Sol, and Opus 5.5 at max effort 2 points higher still. If Claude is on your shortlist, our Sonnet 5.5 benchmarks show where that 4-point lead comes from. For the top of Anthropic's range, the Opus 5.5 benchmarks break down the flagship's full results.

Our read: pick 6.1 Sol when cost per task drives the decision, such as agents that loop many times or workflows with heavy cached context. Pick a Claude model when you need the highest general quality and can absorb the cost. For anything in between, run your own task on both.

What the benchmarks do not tell you

Benchmarks rank models, but they do not predict behavior inside your product. Several gaps apply here:

  • Most scores are vendor-run. OpenAI ran its tests in its own research environment, and it notes results may differ from production ChatGPT.
  • Few absolute numbers. Outside DeepSWE, OpenAI reports gains in points or cost ratios rather than raw scores, which makes cross-checking harder.
  • Cross-vendor costs are not like for like. OpenAI notes its Claude Fable 5.1 cost figure on AutomationBench leaves out fallbacks, which happened on about 40% of tasks.
  • The hard prompts are hard on purpose. The factuality and safety tests use prompts chosen to trigger failures, so real-world error rates should be lower.
  • Independent scores move. Artificial Analysis updates its figures as it retests, so check the live page before a big decision.

GPT-6.1 Sol is the first model to test for cost-conscious agent work

GPT-6.1 Sol benchmarks back up OpenAI's pitch. Independent testing places it 1 point behind GPT-6 Astra at under a quarter of the cost per task, and OpenAI's own DeepSWE data shows it winning outright on coding at high effort.

If you already use GPT-6 Sol, switching is an easy call, as long as you replace any "none" effort setting and check latency. In every case, confirm success rates and tool use on your own tasks before moving production traffic. If you use Astra for everyday agent work, test 6.1 Sol first and keep Astra for hard research. If you are weighing Claude, compare both on your own task, because the vendor and independent results point in different directions.

Build with frontier models without writing code

Picking the model is half the job; you still need the app around it. Emergent turns a plain-language description into a full-stack, deployable app, so founders and operators can put models like GPT-6 Sol to work without hiring engineers.

Emergent connects OpenAI's GPT models, Anthropic's Claude, and Google's Gemini through the Universal LLM Key, with one credential and one bill. You choose the model when you start a project.

Start Building on Emergent.

Was this article helpful?
About the writer

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Cta image

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free
Share this article:

Frequently Asked Questions

Your Questions, Answered

Is GPT-6.1 Sol better than GPT-6 Astra?
Not overall, but it comes close. Artificial Analysis scores GPT-6.1 Sol 1 point below Astra on its Intelligence Index, and OpenAI reports it beating Astra on DeepSWE v1.1 coding at high effort. Astra still leads on hard scientific research, scoring 68.1% on Terminal-Bench Science 0.1. For most agent and coding work, 6.1 Sol delivers similar results at about a quarter of the cost per task.
Is GPT-6.1 Sol cheaper than GPT-6 Sol?
Slightly. Both models list at $2 per 1M input tokens and $10 per 1M output tokens, but GPT-6.1 Sol cuts cache reads from $0.20 to $0.10. Artificial Analysis also measures a lower cost per task at max effort, $0.72 versus $1.05, because 6.1 Sol reaches better results more efficiently. It does use 10% to 30% more output tokens.
What reasoning effort should I use with GPT-6.1 Sol?
Start with high effort for coding and medium for business workflows. On OpenAI's DeepSWE data, GPT-6.1 Sol scores best at high effort (75.2%) and drops to 71.9% at xhigh and max. Artificial Analysis also found xhigh beating max on its coding index. Low effort works well for high-volume tasks. The "none" setting is not supported.
Does GPT-6.1 Sol beat Claude Opus 5.5?
On OpenAI's own tests, yes. GPT-6.1 Sol scores 2.2 points above Opus 5.5 on AutomationBench at medium effort and outscores it on GDP.pdf, both at a lower cost per task. On the independent Artificial Analysis Intelligence Index, Opus 5.5 at max effort scores 58 against 6.1 Sol's 52. The better pick depends on whether cost or peak quality matters more.
Where can I use GPT-6.1 Sol?
GPT-6.1 Sol is available in the OpenAI API as gpt-6.1-sol. It is also in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. It is not yet available in regular ChatGPT chat. OpenAI says an Ultrafast version with up to 8x faster token generation in Codex is coming soon.
Start Building
on Emergent today
Try Emergent

https://api.linear.app/graphql