GPT-6 Sol benchmarks tell a score-per-dollar story, not a peak-score one. OpenAI released GPT-6 Sol on September 22, 2026, about three weeks after flagship GPT-6 Astra, and priced it to make sustained agent work affordable. On standard API rates, Sol is priced at one-fifth of Astra's input and output token rates, and OpenAI positions it as capturing most of Astra's task performance at a fraction of the cost.
That framing matters when you read the numbers. In OpenAI's tests, Sol outscores larger, pricier models on some agentic tasks and trails its own predecessor on others. The through-line is cost efficiency.
This guide covers what Sol scored, what each score cost, how OpenAI reported the figures, and where independent testing agrees or diverges. It closes with how Sol compares to GPT-5.6 Sol, Astra, Luna, and Claude Opus 5.5, plus pricing and specs.
What GPT-6 Sol scored on every benchmark
GPT-6 Sol posts strong agentic scores at a low cost per task. OpenAI evaluated it on six headline tests spanning business workflows, coding, computer use, and factual reliability. The table below shows the best published score for each, with the reasoning effort that produced it and the cost to run one task.
Table 1 - GPT-6 Sol headline benchmark scores (vendor-reported).
Two results anchor OpenAI's pitch. On AutomationBench, Sol at xhigh effort scores 33.2% at $0.27 per task, beating Claude Opus 5 at max effort (26.9%) for about one-eleventh of the cost. On Agents' Last Exam, Sol at max reaches 56.4%, above Opus 5's best score at roughly 60% lower cost per task.
How OpenAI reported these numbers, and why it matters
Read every cross-vendor comparison here as directional, because OpenAI ran the tests two different ways. OpenAI evaluated its own models in its research environment or via its API, then pulled competitor scores from publicly available reports. The models were not run on a single shared harness.
That is standard practice for a launch post, and it does not make the numbers wrong. It does mean a row that places Sol next to a Claude model is comparing two labs' separate runs, not a head-to-head on identical infrastructure. The Claude figures OpenAI cites are for Opus 5 and Fable 5, not the Claude Opus 5.5 that Anthropic launched the same day, so they should not be read as a comparison against Anthropic's current flagship.
OpenAI's launch post is also selective about which comparisons it highlights. It leads with the tests where Sol looks strongest against Anthropic and gives less prominence to the two where Sol trails its own predecessor. Both patterns are worth knowing before you treat any single figure as settled. For a clean same-harness read, independent testing is the better guide.
GPT-6 Sol on the Artificial Analysis Intelligence Index
Independent testing puts GPT-6 Sol's intelligence roughly level with GPT-5.6 Sol, at about half the cost. Artificial Analysis, which runs every model through the same harness, scores Sol at 48 on its Intelligence Index at max effort. The index combines 10 evaluations covering reasoning, knowledge, coding, and long-context work.
Table 2 - GPT-6 Sol on the Artificial Analysis Intelligence Index (independent, same-harness).
The independent read is more measured than OpenAI's framing. In its analysis of the release, Artificial Analysis found that Intelligence Index and Coding Agent Index scores hold level with GPT-5.6, with gains on some evaluations and regressions on others. Sol rises on AutomationBench-AA and Terminal-Bench 4.0 but drops about 100 Elo on GDPval-AA v2.1, a knowledge-work evaluation, which Artificial Analysis attributes to shorter deliverables that more often omit required elements. GPT-6 Sol at max costs $1.06 per Intelligence Index task, about 50% less than GPT-5.6 Sol at $1.99.
Hallucination shows a real gain with a trade-off. On the AA-Omniscience benchmark, Sol at max cut its hallucination rate from 92% to 60%, partly because it declines to answer more often. Its attempted-answer rate fell from 99% to 83%, and its accuracy dropped from 59% to 54%, while its AA-Omniscience Index rose from 22 to 27.
GPT-6 Sol benchmark scores by reasoning effort
Effort changes both the score and the bill, and higher is not always better. GPT-6 Sol runs at six settings, from non-reasoning up to max. The table below shows how each headline benchmark moves across the five reasoning levels.
Table 3 - GPT-6 Sol scores by reasoning effort (vendor-reported).
1. What the effort dial buys you
More effort mostly raises coding and computer-use scores. DeepSWE climbs from 37.2% at low to 68.8% at max, and OSWorld rises from 43.9% at low to 60.5% at xhigh. The cost climbs with it, so higher effort makes sense for hard, high-value tasks rather than routine ones.
2. Where more effort stops helping
Two benchmarks peak below max. Sol scores highest on AutomationBench at xhigh (33.2%) and dips at max (32.0%). Its factual error rate bottoms out at xhigh (4.5%) and ticks up slightly at max. The practical takeaway is to test the effort setting on your own workload instead of defaulting to max.
GPT-6 Sol vs GPT-5.6 Sol: what actually improved
GPT-6 Sol is a clear upgrade on cost and a mixed one on peak score. At max effort it improves on most benchmarks while cost per task falls by roughly half or more. The one clear peak-score regression is DeepSWE, which OpenAI's prose does not call out.
Table 4 - GPT-6 Sol vs GPT-5.6 Sol at max effort (vendor-reported).
DeepSWE is the clear regression, where GPT-6 Sol's best score (68.8%) sits below GPT-5.6 Sol's best (72.7%). Sol gives up a few points of top-end coding performance in exchange for a cost per task that drops roughly 58%.
Independent data supports the direction: on llm-stats, GPT-5.6 Sol ranks above GPT-6 Sol on DeepSWE. At matched cost, GPT-6 Sol comes out ahead, since Sol at high effort beats GPT-5.6 Sol at medium for less than half the price.
OSWorld is not a clear regression. OpenAI highlights a 60.5% result for Sol at xhigh, close to Claude Opus 5 at medium, and independent rankings put GPT-6 Sol level with or slightly above GPT-5.6 Sol.
Factuality improves. OpenAI reports Sol makes about half as many mistakes as GPT-5.6 Sol on its internal factuality evaluation. That evaluation uses error-inducing, user-flagged conversations rather than typical traffic, so read the result as directional rather than a general error rate.
GPT-6 Sol vs Claude Opus 5.5, Astra, and Luna
Sol competes on price rather than peak capability against the strongest current models. Anthropic released Claude Opus 5.5 the same day as Sol, but OpenAI benchmarked against the older Opus 5. Anthropic's own comparison table lists GPT-5.6 Sol, not GPT-6 Sol, so there is no published head-to-head between the two. Placing each vendor's own reported figures side by side, Opus 5.5 reports higher scores than Sol on the benchmarks they share, though these are separate vendor runs rather than a controlled ranking.
Table 5 - GPT-6 Sol against rival models (mixed sources, see caption).
Using each vendor's separately reported figures, Opus 5.5's published score is 6.8 points above Sol's on AutomationBench and 5.1 points above on FrontierCode, and it posts a Terminal-Bench 4.0 result that tops even Astra. Treat those gaps as directional, since the labs did not publish a shared harness. Sol's case is price, at half Opus 5.5's per-token rate and a lower cost per task than the Claude models OpenAI measured.
Against Astra, Sol trades depth for economy. Astra costs 5 times as much per token and leads every chart, most clearly on computer use. For most agent pipelines, Sol is the practical default, with Astra held in reserve for tasks where the extra accuracy justifies the premium.
Luna is the budget story. At $0.10 and $0.50 per million tokens, it scores 66.6% on DeepSWE at $0.22 per task, close to Opus 5 at medium effort for a fraction of the cost.
GPT-6 Sol pricing and what a task actually costs
GPT-6 Sol costs $2 per 1M input tokens and $10 per 1M output tokens, half of GPT-5.6 Sol's rate. Cached input reads carry a 90% discount, which matters most for agents that reread large system prompts and codebases on every turn.
Table 6 - GPT-6 family pricing (as of September 2026).
Three details are easy to miss. Cache writes on Sol are billed at $2.50 per million tokens. Requests over 272,000 input tokens carry a long-context surcharge of 2 times the input rate and 1.5 times the output rate. In the API, Batch and Flex processing run at 50% of standard, while Fast mode runs at 2 times standard.
At $2 and $10 per million tokens, Sol's API rate is half the $4 and $20 that Anthropic lists for Claude Opus 5.5. Per-token price is only part of the picture, since caching, tool calls, output length, and reasoning all shape the final cost per task.
GPT-6 Sol specs and availability
GPT-6 Sol pairs a large context window with the full set of OpenAI agent tools. It handles text and image input, returns text, and runs at six reasoning-effort settings.
Table 7 - GPT-6 Sol specifications.
Availability spans ChatGPT and the API. At launch, OpenAI said Sol was rolling out gradually in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, and was not yet in the standard Chat experience. In the API it is available as gpt-6-sol. Use the Responses API for built-in tools and reasoning; Chat Completions is supported, but function calling there requires reasoning effort set to none.
Which GPT-6 model should you use?
Match the model to the workload, since the three GPT-6 tiers now sit far apart on cost. Luna covers cheap, high-volume tasks. Sol handles everyday agent and automation work. Astra and Claude Opus 5.5 are for peak reasoning.
Table 8 - Choosing across the GPT-6 lineup.
Build with GPT-6 Sol on Emergent
You can build production-grade apps on GPT-6 Sol without touching an API. Emergent is a vibe coding platform where you describe the app you want and its multi-agent architecture builds, tests, and deploys the full stack for you. GPT-6 Sol is available on Emergent today, so its low cost per task translates directly into affordable agentic features in the apps you ship.
The Universal LLM Key makes model choice a one-line decision. It gives you GPT, Claude, and Gemini through a single credential and unified billing, so you can route reasoning-heavy steps to a frontier model and high-volume steps to a cheaper one like Sol. There are no separate accounts and no API keys to juggle.
The benchmark picture points to a clear pattern for builders. Use GPT-6 Sol as the workhorse behind chatbots, automations, and internal tools, and reserve a heavier model for the rare step that needs peak reasoning.
Start Building on Emergent and put GPT-6 Sol to work on your next app.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







