GPT-6.1 Sol pricing looks identical to its predecessor at first glance, and for most line items it is. OpenAI kept the $2 input and $10 output rates when it released GPT-6.1 Sol on September 29, 2026, and changed only one number: the price of reading cached context.
That one number matters more than it sounds. Agents, support bots, and document tools reread the same instructions and reference material on every call, so cache reads can make up most of the tokens you pay for. Halving them cuts the bill for exactly the apps founders are building now.
The rest of the rate card has sharp edges worth knowing before you ship: a long-context cliff, reasoning tokens billed as output, and speed tiers that multiply everything. For the rates this model inherits, see our GPT-6 Sol pricing guide.
GPT-6.1 Sol pricing at a glance
GPT-6.1 Sol costs $2 per 1M input tokens, $0.10 per 1M cache reads, $2.50 per 1M cache writes, and $10 per 1M output tokens on Standard processing. Those rates apply to requests up to 272K input tokens, according to OpenAI's API pricing page.
Table 1 - GPT-6.1 Sol Standard pricing per 1M tokens (OpenAI, pricing as of September 2026).
The API model name is gpt-6.1-sol. Regional processing adds a 10% premium where it is available.
What changed from GPT-6 Sol pricing
Only the cache read rate changed. It fell from $0.20 to $0.10 per 1M tokens, a 50% cut, while input, cache writes, and output stayed exactly where GPT-6 Sol had them.
Table 2 - GPT-6 Sol vs GPT-6.1 Sol Standard rates per 1M tokens, up to 272K input (OpenAI, pricing as of September 2026).
Because OpenAI set the cache discount as a multiplier, the halving carries through every other row too. Long-context, Batch, and Fast cache reads are all half their GPT-6 Sol equivalents.
The upgrade is not free in every case, though. In its launch analysis, Artificial Analysis found GPT-6.1 Sol uses 10% to 30% more output tokens than GPT-6 Sol, and output is the most expensive line on the card. Even so, its measured cost per task at max effort is lower: $0.72 versus $1.05.
How prompt caching changes the bill
Prompt caching is where GPT-6.1 Sol saves real money. A cache read costs 5% of the normal input rate, down from 10% on GPT-6 Sol, per the GPT-6.1 Sol model page.
Caching works like this: the first time you send a long, repeated prefix (system instructions, a product catalog, a policy document), you pay to write it to the cache. Every later call that reuses that prefix pays the cheap read rate instead of full input price.
- Cache writes: Billed at 1.25x the input rate, so $2.50 per 1M tokens. You pay this once per cached prefix.
- Cache reads: Billed at $0.10 per 1M tokens, 95% below standard input.
- Break-even: OpenAI bills written tokens at 1.25x the input rate, so the first call costs $2.50 per 1M prefix tokens instead of $2.00. Each reuse then costs $0.10 instead of $2.00, which recovers that extra $0.50 on the first cache hit. Check your usage logs after launch to confirm the write and read lines match.
The practical takeaway: structure prompts so the stable part comes first and the changing part comes last. An agent that sends a 50K-token knowledge base with every question pays $0.10 for that context at full input price, or half a cent once it is cached.
The 272K long-context rule
Once a request passes 272K input tokens, the entire request moves to long-context pricing, not just the tokens above the line. Input and cache rates double and output rises by 1.5x for the whole call.
That creates a cliff. For input-heavy requests the cost nearly doubles; when output makes up more of the bill, the jump is smaller, because output rises by 50% rather than 100%. Using list rates, a request with 270K input tokens and 5K output tokens costs about $0.59. Add 10K more input tokens, and the same request costs about $1.20.
Table 3 - Estimated cost of one request either side of the 272K threshold, from OpenAI list rates with assumed token counts (pricing as of September 2026).
If your app feeds whole contracts, codebases, or long chat histories into the model, keep an eye on this threshold. Trimming or summarizing context to stay under 272K is often the single cheapest optimization available.
Batch, Flex, Fast, and Ultrafast pricing
OpenAI sells the same model at four speeds, and the price moves with the speed. Batch and Flex cost 50% less than Standard, and Fast mode costs twice as much.
Table 4 - GPT-6.1 Sol pricing by processing tier per 1M tokens, up to 272K input (OpenAI, pricing as of September 2026).
Batch suits work nobody waits on, like overnight report generation or tagging a backlog of support tickets. Fast mode, which OpenAI renamed from Priority processing in July 2026, is for apps where a person is watching the screen.
A faster tier is coming. OpenAI says GPT-6.1 Sol Ultrafast, with up to 8x faster generation in Codex, will arrive in the coming days. OpenAI had not published GPT-6.1 Sol Ultrafast rates as of September 30, 2026. VentureBeat reports that OpenAI's DevDay materials put Ultrafast API usage at 6x Standard. If that holds, GPT-6.1 Sol Ultrafast would cost roughly $12 input and $60 output per 1M tokens. Treat those as reported estimates, not list prices.
What a GPT-6.1 Sol task actually costs
Token prices tell you the rate, not the bill. What you actually pay depends on how many tokens the model spends thinking, and GPT-6.1 Sol bills its reasoning tokens at the output rate.
Artificial Analysis measures this directly by running the same test suite at every effort level. Its release comparison shows cost per task rising more than fivefold from low to max effort.
Table 5 - GPT-6.1 Sol cost per Artificial Analysis Intelligence Index task by reasoning effort (independent, as of September 2026).
The jump from xhigh to max nearly doubles the cost. The extra spend does not always buy better results: Artificial Analysis found xhigh outscored max by 3 points on its Coding Agent Index. Our GPT-6.1 Sol benchmarks breakdown covers how each effort level performs. Start at medium, which is the default, and move up only when your own results justify it.
GPT-6.1 Sol vs Astra, Luna, and Claude pricing
GPT-6.1 Sol sits in the middle of OpenAI's lineup and matches Claude Sonnet 5.5 on list price. Its advantage over Sonnet 5.5 comes down to one line: cache reads at half the price.
Table 6 - Standard list prices per 1M tokens for comparable models (OpenAI and Anthropic, pricing as of September 2026).
Astra costs exactly five times as much for standard input and output, and 10 times as much for cache reads, per OpenAI's launch announcement. Luna is 20 times cheaper than Sol on input and output, and our Sol and Luna explainer covers when the lighter model is enough.
On the Claude side, Anthropic lists Sonnet 5.5 at the same $2 and $10, with $0.20 cache reads and $2.50 cache writes, on its Sonnet 5.5 page. Opus 5.5 costs twice that on input and output. Anthropic also says Sonnet 5.5 needs fewer tokens per task than its predecessor, so identical list prices do not guarantee identical bills. Our Sonnet 5.5 pricing guide has the full Claude rate card.
Estimating a monthly bill
A low-volume, heavily cached app can run for tens of dollars a month on GPT-6.1 Sol. Take a support assistant that answers 1,000 questions a month, rereads a 50K-token knowledge base each time, receives 2K tokens of new input, and writes 3K tokens of output including reasoning.
Table 7 - Estimated monthly cost for 1,000 support conversations, from OpenAI list rates with assumed token counts (pricing as of September 2026).
Astra's figure breaks down as $0.05 for cached context at $1 per 1M, $0.02 for new input, and $0.15 for output. These figures assume every conversation hits the cache and leave out cache writes and tool fees. Bills grow quickly with more volume, longer outputs, higher effort, or frequent cache misses. Still, the ratios hold: 6.1 Sol runs the same workload for about a sixth of Astra's price.
Output is the biggest line here. At 3K output tokens, output accounts for $0.03 of the $0.039 per conversation, which is why lowering reasoning effort usually saves more than any caching trick.
Where you can use GPT-6.1 Sol
GPT-6.1 Sol is available in the OpenAI API, in ChatGPT Work and Codex, and on Microsoft's cloud. OpenAI offers it to Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. It is not yet available in regular ChatGPT chat. The rates in this guide are API token prices; ChatGPT plans follow their own subscription terms, not this rate card.
GPT-6.1 Sol is also generally available in Microsoft Foundry, for teams that already buy AI through Azure. Azure pricing can differ from OpenAI's by deployment type and data zone, so check Microsoft's rates before assuming parity.
GPT-6.1 Sol is the cheapest way to get near-Astra results
GPT-6.1 Sol pricing makes a clear case. You get the same $2 and $10 rates as GPT-6 Sol, half-price cache reads, and a model that independent testing puts within a point of Astra. Artificial Analysis found no cheaper model at its intelligence level.
To keep costs down, cache your stable context, stay under 272K input tokens per request, start at medium effort, and send any work that can wait to Batch. Those four habits matter more than the headline rate.
Build on frontier models and keep the bill predictable
A cheap model still needs an app around it. Emergent turns a plain-language description into a full-stack, deployable app, so founders and operators can build agents and tools without hiring engineers.
Emergent runs OpenAI's GPT models, Anthropic's Claude, and Google's Gemini through the Universal LLM Key, with model usage billed in Emergent Credits instead of separate provider accounts.
Start Building on Emergent.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







