GPT-6.1 Sol arrived seven days after GPT-6 Sol at the same list price. Here is what changed, what it costs in practice, and who should switch.
GPT-6.1 Sol is the stronger model for complex work, and it carries the same list price as GPT-6 Sol. OpenAI released it on September 29, 2026, only seven days after GPT-6 Sol, according to Artificial Analysis.
For anyone who builds software with AI, including the non-technical founders and operators who use Emergent, model choice changes what a build costs and how often it breaks. This guide compares the two models using OpenAI announcements, OpenAI API documentation and independent tests. When a figure comes from OpenAI's own testing, we say so.
What is GPT-6.1 Sol?
GPT-6.1 Sol is OpenAI's balanced model for complex coding, computer use, and professional work. It upgrades GPT-6 Sol and sits between GPT-6 Astra, the flagship, and GPT-6 Luna, the low-cost option. Developers call it with the API model ID gpt-6.1-sol.
OpenAI introduced GPT-6 Sol and Luna on September 22, 2026. It cut Sol's API price by 50% versus GPT-5.6 Sol, from $4 and $20 to $2 and $10 per million input and output tokens.
A token is a small chunk of text, often a short word or part of one.
In its GPT-6.1 Sol announcement, OpenAI says the model nearly matches GPT-6 Astra on agentic coding, computer use, and professional work. It charges one-fifth of Astra's standard input and output prices to do it.
GPT-6.1 Sol is rolling out to eligible paid users in ChatGPT Work and Codex, with access depending on plan and workspace settings. It is not available in regular ChatGPT conversations. Developers can access it through the API.
GPT-6.1 Sol vs GPT-6 Sol at a glance
The specs look nearly identical. The differences sit in capability and in how you call the model.
Pricing as of October 2026. Sources: OpenAI announcements and API documentation, plus the Intelligence Index and context window figures on OpenRouter's comparison page.
What stays the same between the two models
GPT-6.1 Sol changes how well the model works, not the shape of the product. Both models take text and image inputs and return text. Both offer a 1,050,000-token context window, so a large codebase or a stack of documents fits in one request.
Both default to medium reasoning effort, and both use the same $2 and $10 standard rates. OpenAI's model page lists an April 30, 2026 knowledge cutoff for GPT-6.1 Sol. OpenAI lists an April 30, 2026 knowledge cutoff for GPT-6.1 Sol, so use search or file tools when the task depends on newer information.
For most teams, this means the swap does not change what you can send or receive. It changes how well the model handles the work, and how you configure the call.
Staying within OpenAI isn't the only option. Our GPT-6.1 Sol alternatives guide covers what else competes at this tier.
What changed in GPT-6.1 Sol
GPT-6.1 Sol improves on GPT-6 Sol in all five areas OpenAI tested. Every figure in this section comes from OpenAI's own evaluations, and OpenAI took competitor scores from public reports.
1. Coding
GPT-6.1 Sol beats GPT-6 Sol's best DeepSWE v1.1 score by 6.4 percentage points. DeepSWE v1.1 tests complex software-engineering tasks in real codebases. OpenAI says the new model also matches Astra on this test at roughly one-fifth of the cost.
The 6.4-point gain came at a lower reasoning effort than GPT-6 Sol's best run. Reasoning effort is a setting that controls how long the model thinks before it answers. Because the settings differ, read the gap as a direction, not an exact margin.
2. Business workflows
On AutomationBench 1.0.6, GPT-6.1 Sol scores 4.8 percentage points above GPT-6 Sol at medium effort. The test runs agents through end-to-end workflows using 47 tools across sales, marketing, operations, support, finance, and HR. OpenAI also reports a 2.2-point lead over Claude Opus 5.5 at roughly one-third of the cost.
On GDP.pdf, which asks professional questions about complex PDFs, OpenAI says GPT-6.1 Sol outscores Opus 5.5 with fallbacks at less than half the cost per task.
3. Computer use
Computer use means the model operates apps by clicking and typing, as a person would. On the offline set of OSWorld 2.0, GPT-6.1 Sol beats GPT-6 Sol by seven percentage points at max effort, at less than half the cost. It lands within 2.1 points of Astra at roughly one-seventh of Astra's cost per task.
4. Scientific research
On Terminal-Bench Science 0.1, GPT-6.1 Sol more than doubles GPT-6 Sol's max-effort score. It costs $5.47 per task on average, against $23.21 for Opus 5.5 and $23.80 for Astra. Astra still posts the top score at 68.1%, and OpenAI says to use it for the hardest research tasks.
5. Factual accuracy
At low reasoning effort, the share of responses containing a factual error falls from 11.4% to 7.7%, a reduction of 3.7 percentage points, or about 32% relative to GPT-6 Sol. Across the tested settings, GPT-6.1 Sol stays within 1.9 percentage points of Astra's error rate.
OpenAI built this test from ChatGPT conversations where users had flagged earlier mistakes. The prompts are deliberately hard and do not represent typical use.
What independent testing shows
Artificial Analysis's independent testing points in the same direction as OpenAI's results. On its Intelligence Index, GPT-6.1 Sol gains four points over GPT-6 Sol and five over GPT-5.6 Sol. It sits one point below GPT-6 Astra.
The cost data matters more for builders. At max effort, GPT-6.1 Sol costs $0.72 per Intelligence Index task, against $1.05 for GPT-6 Sol and $3.26 for Astra. That is 31% less than its predecessor and under one-quarter of Astra's price.
The firm also reports gains on individual tests: 12 points on Terminal-Bench 4.0, five on Humanity's Last Exam, and six on GDP.pdf. On AA-Omniscience, the hallucination rate falls from 60% to 54%. In its Coding Agent Index, GPT-6.1 Sol scores one point above Astra at high effort for under 15% of the cost per task.
One caveat applies. GPT-6.1 Sol uses 10-30% more output tokens than GPT-6 Sol at the same effort setting. Artificial Analysis still measured a lower cost per task at max effort.
How GPT-6.1 Sol handles its own limits
OpenAI reports that GPT-6.1 Sol is more honest about what it cannot do. In one test, an agent's search tool is broken, and the question is whether the model tells the user or guesses. GPT-6.1 Sol fails to disclose the problem in 2.1% of cases, against 4.9% for GPT-6 Sol and 1.5% for Astra. GPT-6 Luna fails in 28.7% of cases.
That gap matters for anyone who lets an agent run unattended. A model that hides a broken tool hands you a confident wrong answer.
OpenAI also says it saw no attempts by GPT-6.1 Sol to bypass an automated safety reviewer, matching Astra and GPT-6 Sol. It reports lower failure rates than GPT-6 Sol at respecting explicit restrictions and avoiding unauthorized outcomes. OpenAI notes that these tests are chosen to provoke failures, so they do not reflect typical use. The full results sit in OpenAI's system card addendum.
GPT-6.1 Sol pricing: where the savings are
GPT-6.1 Sol costs the same as GPT-6 Sol on fresh input and output, and half as much on cached input. Cached input is text the model already processed in an earlier request, such as a long project brief or a codebase. Agents resend that text at every step, so the discount adds up.
Pricing as of October 2026. Prices are per 1 million tokens at standard short-context API rates. Source: OpenAI.
Here is a worked example. Say an agent reuses a 200,000-token project context across 50 requests. That is 10 million cached tokens. At GPT-6 Sol's $0.20 rate, the reads cost $2.00. At GPT-6.1 Sol's $0.10 rate, they cost $1.00. This is our own arithmetic from OpenAI's published rates, so treat it as an illustration.
Two limits apply. For prompts above 272,000 input tokens, OpenAI charges 2× the standard input and cached-input rates and 1.5× the standard output rate across the full request. And OpenAI plans a GPT-6.1 Sol Ultrafast tier, which VentureBeat reports costs six times the standard rate and reaches up to 300 tokens per second.
For the cross-vendor comparison, read our Sonnet 5.5 vs GPT-6.1 Sol breakdown.
What you give up by switching
GPT-6.1 Sol brings three trade-offs, and two of them require code changes.
1. No, none or minimal reasoning
GPT-6.1 Sol always reasons. OpenAI's model page says effort supports low, medium (the default), high, xhigh, and max. The none and minimal settings are not supported.
If your app used none for quick, cheap replies, OpenAI's guidance is to use low instead. Check latency after the swap.
2. Tool calling needs the Responses API
A tool call lets the model trigger an action, such as a web search or a database update. With GPT-6.1 Sol, tool calling requires OpenAI's Responses API. Chat Completions still works, but without tools.
GPT-6 Sol allowed function calling in Chat Completions only when effort was set to none. Apps built that way must move. OpenAI also says to remove temperature, top_p, and top_logprobs whenever reasoning is on, which is always the case here.
3. Slower output
Independent measurements currently show GPT-6.1 Sol generating output more slowly than GPT-6 Sol. Exact throughput varies by provider, reasoning effort, and testing conditions, so benchmark the difference on your own workload if latency matters.
You pay for deeper thinking with time.
How to migrate from GPT-6 Sol to GPT-6.1 Sol
A safe migration takes six steps, and most take minutes. The table follows OpenAI's migration guidance.
Source: Source for Steps 1–5: OpenAI API documentation. Step 6 is a suggested validation approach for comparing the models on your own workload.
Your own test beats any benchmark. Pick tasks from real work, such as a bug fix, a report summary, or a form-filling flow. Run each on both models at the same effort. A four-column sheet is enough: task, pass or fail, cost, and seconds.
A one-afternoon test plan
You can compare the two models in a few hours, and the results rest on your data. Follow these six steps.
- Choose 20 tasks from real work. Mix easy and hard ones, such as a bug fix, a document summary, and a multi-step workflow.
- Run every task on both models at the same reasoning effort. Start at medium, since both default to it.
- Score each result as pass or fail. Use the same person and the same checklist for both models.
- Record cost and time per task. Pull token usage from the API response and calculate the cost using the applicable input, cached-input, and output rates.
- Repeat the hardest five tasks at low and high effort. This shows where extra thinking pays off and where it only adds cost.
- Compare the totals. Pick GPT-6.1 Sol if it passes more tasks at a cost and speed you accept.
If your numbers disagree with a leaderboard, trust your numbers. They reflect your prompts, your data, and how long you are willing to wait.
Who should switch now, and who should wait
Most teams using GPT-6 Sol should test GPT-6.1 Sol, especially for new projects. The list price is unchanged, cached reads are cheaper, and independent testing shows a real capability gain.
Switch now if:
- Your agent writes or fixes code in real codebases.
- Your agent reuses a long context across many steps, so the cached-input discount applies.
- Your workflow spans several apps, or the model operates a computer.
- You want near-Astra results without Astra's $10 and $50 prices.
Wait, or pick a different model, if:
- Your app calls tools through Chat Completions with none. Move to Responses first, or stay on GPT-6 Sol.
- Speed matters more than depth. OpenAI positions GPT-6 Luna as the low-cost choice for focused, high-volume work.
- You run the hardest scientific research tasks. OpenAI says Astra still leads there.
What this means if you build apps with AI
A better model helps app builders only when it lifts real work, not just benchmark scores. The skills measured here map to what an AI app builder does all day: editing real code, using tools, and finishing multi-step jobs.
Platforms such as Emergent use AI to help users build full-stack applications with components such as backends, databases, authentication, and deployment. Improvements in coding, tool use, and multi-step workflows can therefore matter to the quality and reliability of the building process.
Do not read a benchmark as a promise about your app. OpenAI says its factuality and safety tests are harder than typical use. Run your own workflow, then decide.
Test one workload on GPT-6.1 Sol this week.
GPT-6.1 Sol earns a test slot on any agent that writes code, completes workflows, or reuses long context. The list price is flat, the cached rate is half, and independent testing shows a real capability gain.
Start small. Pick one workload, run it on both models, and compare cost, quality, and time. Fix the none setting and the tool-calling path first, because those two changes break code.
If you would rather describe the business you want to run than tune a model, visit Emergent

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







