HomeLearn

GPT-6 Sol Benchmarks: Scores, Cost, and How It Compares

GPT-6 Sol benchmarks explained: scores on 6 agentic tests, cost per task, full pricing, and how Sol compares with GPT-5.6 Sol, Astra, and Claude Opus 5.5.

Bhavyadeep
Written by
Bhavyadeep
Saurabh Anand
Reviewed by
Saurabh Anand
Last updated: 
September 23, 2026
0
 min read
Select Emergent as your Preferred news source
Table of Contents

TL;DR

  • GPT-6 Sol is a cost-efficiency release. OpenAI keeps Astra as its flagship and positions Sol to deliver most of that capability at a fraction of the cost.
  • On OpenAI's own tests, Sol scores 33.2% on AutomationBench, 56.4% on Agents' Last Exam, 49.3% on FrontierCode 1.1, and 68.8% on DeepSWE v1.1.
  • Independently, Artificial Analysis puts Sol at 48 on its Intelligence Index at max effort, level with GPT-5.6 Sol on intelligence while cost per task roughly halves.
  • Pricing dropped 50% to $2 per 1M input tokens and $10 per 1M output tokens (as of September 2026).
  • Sol is a strong default for agentic and business-automation work; validate it on your own tasks before standardizing. Step up to Astra or Claude Opus 5.5 for peak coding and reasoning.


GPT-6 Sol benchmarks tell a score-per-dollar story, not a peak-score one. OpenAI released GPT-6 Sol on September 22, 2026, about three weeks after flagship GPT-6 Astra, and priced it to make sustained agent work affordable. On standard API rates, Sol is priced at one-fifth of Astra's input and output token rates, and OpenAI positions it as capturing most of Astra's task performance at a fraction of the cost.

That framing matters when you read the numbers. In OpenAI's tests, Sol outscores larger, pricier models on some agentic tasks and trails its own predecessor on others. The through-line is cost efficiency.

This guide covers what Sol scored, what each score cost, how OpenAI reported the figures, and where independent testing agrees or diverges. It closes with how Sol compares to GPT-5.6 Sol, Astra, Luna, and Claude Opus 5.5, plus pricing and specs.

What GPT-6 Sol scored on every benchmark

GPT-6 Sol posts strong agentic scores at a low cost per task. OpenAI evaluated it on six headline tests spanning business workflows, coding, computer use, and factual reliability. The table below shows the best published score for each, with the reasoning effort that produced it and the cost to run one task.

Table 1 - GPT-6 Sol headline benchmark scores (vendor-reported).

Benchmark What it measures GPT-6 Sol (best effort) Cost per task
AutomationBench 1.0.6 Business workflows across 47 tools 33.2% (xhigh) $0.27
Agents' Last Exam V1 Long-horizon professional tasks 56.4% (max) $2.93
FrontierCode 1.1 Main Mergeable agent-written code 49.3% (max) $2.14
DeepSWE v1.1 Repository-level software fixes 68.8% (max) $2.74
OSWorld 2.0 (offline) Computer-use workflows 60.5% (xhigh) $2.21
Factual error rate Mistakes on error-prone prompts (lower is better) 4.5% (xhigh) $0.13

Two results anchor OpenAI's pitch. On AutomationBench, Sol at xhigh effort scores 33.2% at $0.27 per task, beating Claude Opus 5 at max effort (26.9%) for about one-eleventh of the cost. On Agents' Last Exam, Sol at max reaches 56.4%, above Opus 5's best score at roughly 60% lower cost per task.

How OpenAI reported these numbers, and why it matters

Read every cross-vendor comparison here as directional, because OpenAI ran the tests two different ways. OpenAI evaluated its own models in its research environment or via its API, then pulled competitor scores from publicly available reports. The models were not run on a single shared harness.

That is standard practice for a launch post, and it does not make the numbers wrong. It does mean a row that places Sol next to a Claude model is comparing two labs' separate runs, not a head-to-head on identical infrastructure. The Claude figures OpenAI cites are for Opus 5 and Fable 5, not the Claude Opus 5.5 that Anthropic launched the same day, so they should not be read as a comparison against Anthropic's current flagship.

OpenAI's launch post is also selective about which comparisons it highlights. It leads with the tests where Sol looks strongest against Anthropic and gives less prominence to the two where Sol trails its own predecessor. Both patterns are worth knowing before you treat any single figure as settled. For a clean same-harness read, independent testing is the better guide.

GPT-6 Sol on the Artificial Analysis Intelligence Index

Independent testing puts GPT-6 Sol's intelligence roughly level with GPT-5.6 Sol, at about half the cost. Artificial Analysis, which runs every model through the same harness, scores Sol at 48 on its Intelligence Index at max effort. The index combines 10 evaluations covering reasoning, knowledge, coding, and long-context work.

Table 2 - GPT-6 Sol on the Artificial Analysis Intelligence Index (independent, same-harness).

Reasoning effort Intelligence Index Cost per Intelligence Index task
Max 48 $1.06
xhigh 44 $0.53
High 43 $0.37
Medium 40 $0.25
Low 34 $0.13
Non-reasoning 28 $0.33

The independent read is more measured than OpenAI's framing. In its analysis of the release, Artificial Analysis found that Intelligence Index and Coding Agent Index scores hold level with GPT-5.6, with gains on some evaluations and regressions on others. Sol rises on AutomationBench-AA and Terminal-Bench 4.0 but drops about 100 Elo on GDPval-AA v2.1, a knowledge-work evaluation, which Artificial Analysis attributes to shorter deliverables that more often omit required elements. GPT-6 Sol at max costs $1.06 per Intelligence Index task, about 50% less than GPT-5.6 Sol at $1.99.

Hallucination shows a real gain with a trade-off. On the AA-Omniscience benchmark, Sol at max cut its hallucination rate from 92% to 60%, partly because it declines to answer more often. Its attempted-answer rate fell from 99% to 83%, and its accuracy dropped from 59% to 54%, while its AA-Omniscience Index rose from 22 to 27.

GPT-6 Sol benchmark scores by reasoning effort

Effort changes both the score and the bill, and higher is not always better. GPT-6 Sol runs at six settings, from non-reasoning up to max. The table below shows how each headline benchmark moves across the five reasoning levels.

Table 3 - GPT-6 Sol scores by reasoning effort (vendor-reported).

Benchmark Low Medium High xhigh Max
AutomationBench 1.0.6 21.2% 26.9% 31.2% 33.2% 32.0%
Agents' Last Exam V1 48.7% 53.1% 52.6% 55.4% 56.4%
FrontierCode 1.1 Main 37.3% 45.9% 47.7% 48.4% 49.3%
DeepSWE v1.1 37.2% 56.6% 65.3% 66.6% 68.8%
OSWorld 2.0 (offline) 43.9% 54.0% 58.3% 60.5% 64.4%
Factual error rate (lower is better) 11.4% 6.9% 5.1% 4.5% 4.6%

1. What the effort dial buys you

More effort mostly raises coding and computer-use scores. DeepSWE climbs from 37.2% at low to 68.8% at max, and OSWorld rises from 43.9% at low to 60.5% at xhigh. The cost climbs with it, so higher effort makes sense for hard, high-value tasks rather than routine ones.

2. Where more effort stops helping

Two benchmarks peak below max. Sol scores highest on AutomationBench at xhigh (33.2%) and dips at max (32.0%). Its factual error rate bottoms out at xhigh (4.5%) and ticks up slightly at max. The practical takeaway is to test the effort setting on your own workload instead of defaulting to max.

GPT-6 Sol vs GPT-5.6 Sol: what actually improved

GPT-6 Sol is a clear upgrade on cost and a mixed one on peak score. At max effort it improves on most benchmarks while cost per task falls by roughly half or more. The one clear peak-score regression is DeepSWE, which OpenAI's prose does not call out.

Table 4 - GPT-6 Sol vs GPT-5.6 Sol at max effort (vendor-reported).

Benchmark (max effort) GPT-5.6 Sol GPT-6 Sol Cost per task change
AutomationBench 1.0.6 28.8% 32.0% (33.2% at xhigh) -49%
Agents' Last Exam V1 52.8% 56.4% -59%
FrontierCode 1.1 Main 47.5% 49.3% -59%
DeepSWE v1.1 72.7% 68.8% -58%
Factual error rate (lower is better) 8.5% 4.6% -79%

DeepSWE is the clear regression, where GPT-6 Sol's best score (68.8%) sits below GPT-5.6 Sol's best (72.7%). Sol gives up a few points of top-end coding performance in exchange for a cost per task that drops roughly 58%.

Independent data supports the direction: on llm-stats, GPT-5.6 Sol ranks above GPT-6 Sol on DeepSWE. At matched cost, GPT-6 Sol comes out ahead, since Sol at high effort beats GPT-5.6 Sol at medium for less than half the price.

OSWorld is not a clear regression. OpenAI highlights a 60.5% result for Sol at xhigh, close to Claude Opus 5 at medium, and independent rankings put GPT-6 Sol level with or slightly above GPT-5.6 Sol.

Factuality improves. OpenAI reports Sol makes about half as many mistakes as GPT-5.6 Sol on its internal factuality evaluation. That evaluation uses error-inducing, user-flagged conversations rather than typical traffic, so read the result as directional rather than a general error rate.

GPT-6 Sol vs Claude Opus 5.5, Astra, and Luna

Sol competes on price rather than peak capability against the strongest current models. Anthropic released Claude Opus 5.5 the same day as Sol, but OpenAI benchmarked against the older Opus 5. Anthropic's own comparison table lists GPT-5.6 Sol, not GPT-6 Sol, so there is no published head-to-head between the two. Placing each vendor's own reported figures side by side, Opus 5.5 reports higher scores than Sol on the benchmarks they share, though these are separate vendor runs rather than a controlled ranking.

Table 5 - GPT-6 Sol against rival models (mixed sources, see caption).

Benchmark GPT-6 Sol GPT-6 Luna Claude Opus 5.5 GPT-6 Astra
AutomationBench 1.0.6 33.2% 20.7% 40.0% 41.4%
FrontierCode 1.1 Main 49.3% 42.4% 54.4% 53.3%
Terminal-Bench 4.0 Not reported Not reported 66.4% 57.9%
Input / output (per 1M) $2 / $10 $0.10 / $0.50 $4 / $20 $10 / $50

Using each vendor's separately reported figures, Opus 5.5's published score is 6.8 points above Sol's on AutomationBench and 5.1 points above on FrontierCode, and it posts a Terminal-Bench 4.0 result that tops even Astra. Treat those gaps as directional, since the labs did not publish a shared harness. Sol's case is price, at half Opus 5.5's per-token rate and a lower cost per task than the Claude models OpenAI measured.

Against Astra, Sol trades depth for economy. Astra costs 5 times as much per token and leads every chart, most clearly on computer use. For most agent pipelines, Sol is the practical default, with Astra held in reserve for tasks where the extra accuracy justifies the premium.

Luna is the budget story. At $0.10 and $0.50 per million tokens, it scores 66.6% on DeepSWE at $0.22 per task, close to Opus 5 at medium effort for a fraction of the cost.

GPT-6 Sol pricing and what a task actually costs

GPT-6 Sol costs $2 per 1M input tokens and $10 per 1M output tokens, half of GPT-5.6 Sol's rate. Cached input reads carry a 90% discount, which matters most for agents that reread large system prompts and codebases on every turn.

Table 6 - GPT-6 family pricing (as of September 2026).

Model Input / 1M Cached input / 1M Output / 1M Change vs GPT-5.6
GPT-6 Sol $2.00 $0.20 $10.00 Input and output down 50%
GPT-6 Luna $0.10 $0.01 $0.50 Input down 50%, output down 58%
GPT-6 Astra $10.00 $1.00 $50.00 New top tier

Three details are easy to miss. Cache writes on Sol are billed at $2.50 per million tokens. Requests over 272,000 input tokens carry a long-context surcharge of 2 times the input rate and 1.5 times the output rate. In the API, Batch and Flex processing run at 50% of standard, while Fast mode runs at 2 times standard.

At $2 and $10 per million tokens, Sol's API rate is half the $4 and $20 that Anthropic lists for Claude Opus 5.5. Per-token price is only part of the picture, since caching, tool calls, output length, and reasoning all shape the final cost per task.

GPT-6 Sol specs and availability

GPT-6 Sol pairs a large context window with the full set of OpenAI agent tools. It handles text and image input, returns text, and runs at six reasoning-effort settings.

Table 7 - GPT-6 Sol specifications.

Spec GPT-6 Sol
Release date September 22, 2026
Context window 1,050,000 tokens
Max output 128,000 tokens
Knowledge cutoff April 20, 2026
Input modalities Text, image
Output modalities Text
Reasoning effort levels none, low, medium, high, xhigh, max
API model string gpt-6-sol

Availability spans ChatGPT and the API. At launch, OpenAI said Sol was rolling out gradually in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, and was not yet in the standard Chat experience. In the API it is available as gpt-6-sol. Use the Responses API for built-in tools and reasoning; Chat Completions is supported, but function calling there requires reasoning effort set to none.

Which GPT-6 model should you use?

Match the model to the workload, since the three GPT-6 tiers now sit far apart on cost. Luna covers cheap, high-volume tasks. Sol handles everyday agent and automation work. Astra and Claude Opus 5.5 are for peak reasoning.

Table 8 - Choosing across the GPT-6 lineup.

If you need Pick Why
High-volume classification, extraction, or routing GPT-6 Luna $0.10 / $0.50 pricing and better factuality than GPT-5.6 Luna
A default model for agents and business automation GPT-6 Sol Tops Opus 5 on AutomationBench in OpenAI's tests, at a much lower cost per task
Peak coding and hard reasoning GPT-6 Astra or Claude Opus 5.5 Astra leads on computer use; Opus 5.5 leads FrontierCode and Terminal-Bench 4.0

Build with GPT-6 Sol on Emergent

You can build production-grade apps on GPT-6 Sol without touching an API. Emergent is a vibe coding platform where you describe the app you want and its multi-agent architecture builds, tests, and deploys the full stack for you. GPT-6 Sol is available on Emergent today, so its low cost per task translates directly into affordable agentic features in the apps you ship.

The Universal LLM Key makes model choice a one-line decision. It gives you GPT, Claude, and Gemini through a single credential and unified billing, so you can route reasoning-heavy steps to a frontier model and high-volume steps to a cheaper one like Sol. There are no separate accounts and no API keys to juggle.

The benchmark picture points to a clear pattern for builders. Use GPT-6 Sol as the workhorse behind chatbots, automations, and internal tools, and reserve a heavier model for the rare step that needs peak reasoning.

Start Building on Emergent and put GPT-6 Sol to work on your next app.

Was this article helpful?
About the writer

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Cta image

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free
Share this article:

Frequently Asked Questions

Your Questions, Answered

When was GPT-6 Sol released?
OpenAI released GPT-6 Sol on September 22, 2026, alongside GPT-6 Luna. It arrived about three weeks after flagship GPT-6 Astra. Sol launched in ChatGPT Work and Codex for paid tiers, and in the API as gpt-6-sol, with a gradual rollout on launch day.
How much does GPT-6 Sol cost?
GPT-6 Sol costs $2 per 1M input tokens and $10 per 1M output tokens, as of September 2026. That is a 50% cut from GPT-5.6 Sol. Cached input reads are discounted 90% to $0.20 per million. Prompts over 272,000 input tokens carry a long-context surcharge.
Who created GPT-6 Sol?
OpenAI created GPT-6 Sol. It sits in the middle of the GPT-6 family, between the flagship GPT-6 Astra and the budget GPT-6 Luna. OpenAI positions Sol as its model for complex coding and agentic workflows that need strong performance at a lower cost than Astra.
What is GPT-6 Sol's knowledge cutoff?
GPT-6 Sol has a knowledge cutoff of April 20, 2026. For anything after that date, the model relies on tools such as web search rather than its training data. Sol supports OpenAI's core agent tools, including function calling, web search, file search, and computer use.
Is GPT-6 Sol multimodal?
Yes, in part. GPT-6 Sol accepts both text and image input, and it returns text output. It does not generate images or video itself. The image input supports vision tasks, so the model can read screenshots, documents, and diagrams as part of an agentic workflow.
What is GPT-6 Sol's context window?
GPT-6 Sol has a context window of 1,050,000 tokens, with up to 128,000 tokens of output. That is large enough to hold long documents, extended conversations, or sizable codebases in a single request. Requests above 272,000 input tokens are billed at a higher long-context rate.
Is GPT-6 Sol better than GPT-5.6 Sol?
It depends on what you measure. GPT-6 Sol improves on factuality and most agentic tests while cutting cost per task by roughly half. Its peak DeepSWE score is slightly lower than GPT-5.6 Sol's. For score per dollar, GPT-6 Sol is the clear upgrade.
How does GPT-6 Sol compare to Claude Opus 5.5?
Each vendor's own figures put Opus 5.5 above Sol on AutomationBench and FrontierCode 1.1, but there is no shared-harness head-to-head. Anthropic's comparison table lists GPT-5.6 Sol, not GPT-6 Sol, and OpenAI benchmarked against the older Opus 5. Sol's clear advantage is price, at half Opus 5.5's per-token cost.
Start Building
on Emergent today
Try Emergent

https://api.linear.app/graphql