OpenAI released GPT-6 Luna on September 22, 2026, and the first question builders asked was whether the cheapest model in the lineup could carry real work. The GPT-6 Luna benchmarks answer that with a clear trade-off. Luna is built for high-volume, latency-sensitive tasks, and it holds its ground on cost per task rather than raw intelligence. If you are weighing Luna for production, the useful question is not whether it is the smartest option. It is whether the numbers justify running it at scale. Here is what the verified data shows, and what OpenAI left out.
What is GPT-6 Luna?
GPT-6 Luna is the lowest-cost model in OpenAI's GPT-6 family, positioned below GPT-6 Sol and the flagship GPT-6 Astra. It targets high-volume tasks such as chat, classification, and lightweight agent work, and at higher reasoning effort it can take on heavier coding and computer-use jobs. It sits in the same family as the flagship, whose GPT-6 Astra benchmarks we covered separately.
Luna arrived alongside GPT-6 Sol in a single launch. You can read the full GPT-6 Luna launch coverage for the announcement details. In the API, Luna ships as gpt-6-luna with six reasoning-effort levels: none, low, medium, high, xhigh, and max. It carries a 1,050,000-token context window and a maximum output of 128,000 tokens, according to OpenAI's API documentation.
GPT-6 Luna benchmark scores at a glance
The table below separates what each number is and where it comes from. That distinction matters here, because in the launch post's body text OpenAI stated only one absolute Luna score and expressed the rest as relative gains. Independent scores come from Artificial Analysis, which tests models on its own harness.
GPT-6 Luna benchmark data as of September 2026, labeled by source. Vendor-reported figures come from OpenAI; independent figures come from Artificial Analysis.
The pattern is consistent. Where a claim carries a hard percentage, it holds up in OpenAI's text. Where OpenAI compared Luna to older models, it often gave a gain or a cost ratio rather than a standalone score.
Intelligence Index and what "level with GPT-5.6" means
Luna's headline independent score is an Intelligence Index of 37 at max effort. That figure drops in a clean line as you lower reasoning effort, reaching 18 with reasoning off. The Index is a blended measure of reasoning, knowledge, and problem-solving that Artificial Analysis runs the same way across models, which makes it a fair point of comparison.
The more important finding is what did not change, and where it did. Artificial Analysis reports that Luna's aggregate Intelligence Index holds roughly level with GPT-5.6, with progress on some evaluations and regressions on others. The coding picture is weaker: Luna's Coding Agent Index came in at 41, down 2 points from GPT-5.6 Luna's 43, measured in OpenAI's Codex harness. So the intelligence tier held steady while the price fell, and coding slipped slightly. That mixed result is the real story of this release, and it explains why some early users felt the new model was a sidegrade rather than an upgrade.
Coding and agentic benchmarks
Luna is priced for tasks that run often and run cheap, so its coding and agent results matter most when you multiply them across thousands of calls. Three tests frame that picture.
1. DeepSWE v1.1
DeepSWE v1.1 tests whether an agent can resolve real, multi-file software issues in a working codebase. Luna scores 66.6% at max effort. OpenAI's benchmark report describes it as comparable to Claude Opus 5 and Claude Fable 5 at medium effort, with Luna costing 93% less per task than Opus 5 and 96% less than Fable 5. This is the single cleanest Luna number in the launch, and it is the one figure most competitors agree on.
One pattern worth keeping in mind: on multi-step coding harnesses, spending more test-time compute on a cheaper model can sometimes close the gap with a heavier model at a lower setting. That makes max-effort Luna worth testing on your own tasks before you rule it out on price alone.
2. AutomationBench 1.0.6
AutomationBench, built by Zapier, tests agents on end-to-end business workflows across 47 tools spanning sales, support, finance, and other functions. Here OpenAI gave a relative result: Luna at high effort improves on GPT-5.6 Luna by 5.4 percentage points while cutting cost per task by 58%. OpenAI's absolute-score table for this test lists Sol, Astra, Opus 5, and Fable 5.1, but not Luna, so a precise Luna figure is not part of the public record.
3. OSWorld 2.0
OSWorld 2.0 measures computer use, meaning an agent driving a real desktop through clicks and windows. OpenAI states that Luna at max effort exceeds GPT-5.6 Sol at medium effort, at roughly one-tenth the cost per task. It did not publish a standalone Luna percentage in the body text, so treat any exact OSWorld number for Luna on other sites with care, and check whether it was independently measured or pulled from a vendor chart.
Speed and cost per task
Cost is where Luna makes its case. It runs at $0.10 per 1M input tokens and $0.50 per 1M output tokens, down from $0.20 and $1.20 on GPT-5.6 Luna. Cached input reads are billed at $0.01 per 1M, a 90% discount, which adds up fast for agents that reuse large system prompts on every turn.
Artificial Analysis measures the practical effect of that pricing. Running its full Intelligence Index costs about $0.07 per task on Luna at max effort, roughly 60% less than GPT-5.6 Luna at the same setting. Output speed lands at 157 tokens per second at max effort and ranges from about 140 to 152 across lower settings, which keeps Luna responsive for interactive use.
For a sense of where Luna sits against other budget tiers, the Gemini 3.8 Flash benchmarks make a useful side read, since Flash competes in the same low-cost band. Luna's raw token rates are aggressive for its tier, though the right pick depends on your workload, context sizes, caching terms, and reasoning-effort needs. If you want the fuller pricing picture across the family, the GPT-6 Astra pricing breakdown covers the flagship side.
What OpenAI did not publish
Reading the GPT-6 Luna benchmarks honestly means naming the gaps. OpenAI led its launch with Sol's numbers and gave Luna far less absolute coverage. Three things are missing from the public record:
- Absolute scores on most tests. Outside DeepSWE, Luna's results appear as gains over a predecessor or as cost ratios, not standalone percentages you can drop into a table.
- FrontierCode and Agents' Last Exam. OpenAI reported Sol figures on both, but not Luna.
- Chart-only detail. Some numbers sit inside the launch post's bar charts rather than its text. Precise Luna figures on other sites may come from those charts or from unsourced aggregators, so they carry less certainty than a stated number.
None of this makes Luna a weak model. It does mean a "complete Luna scorecard" is not something the current sources support, and any page that presents one is filling gaps OpenAI left open.
How the benchmarks compare with early user reports
Benchmark scores and daily experience do not always match, and Luna is a case in point. Within hours of launch, some Reddit threads reported that Luna felt weaker than GPT-5.6 Luna on practical tasks. These are unverified, self-selected posts rather than a measured sample, so treat them as sentiment, not evidence. They are worth noting mainly because they rhyme with the one hard independent regression on record: Luna's lower Coding Agent Index.
The takeaway is not that Luna is broken. It is that Luna was tuned to deliver similar intelligence at half the cost, so anyone expecting a step up in raw capability will feel let down, while anyone optimizing for spend at volume gets exactly what the release promised. For coding-heavy pipelines, comparing Luna against a peer like the Grok 4.7 benchmarks can help you calibrate before committing.
Build with GPT, Claude, or Gemini on Emergent
Benchmarks help you pick a model, but the harder job is turning that model into a working product. Emergent lets non-technical founders and operators build production-grade full-stack applications by describing what they want, using leading models from OpenAI, Anthropic, and Google through a single Universal LLM Key. You get to build with frontier AI without wiring up separate accounts or billing.
When you are ready to turn a benchmark decision into a live app, start building on Emergent, the AI app builder handles the full stack from a single description.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







