Claude Sonnet 5.5 is the closest a Sonnet model has come to Opus quality. Anthropic's launch table has it within 3.2 percentage points of Opus 5.5 on every percentage-scored test it trails, and ahead on agentic terminal coding. The catch is in the cost column.
If you are choosing a model for an app, a support workflow, or an internal tool, the headline Sonnet 5.5 benchmarks tell half the story. The effort setting you pick can move the price of one task from under $0.50 to over $7.
Independent testing also disagrees with the launch post on cost, and that gap decides whether Sonnet 5.5 beats the Opus 5.5 benchmarks on value per dollar.
Sonnet 5.5 benchmarks trail Opus 5.5 on 7 of 8 tests, but never by much
Sonnet 5.5 scores close to Opus 5.5 across coding, knowledge work, reasoning, computer use, and chart reading. Anthropic's launch post shades the top score in each row, and Sonnet 5.5 holds that spot only on Terminal-Bench 4.0.
The bigger story is the jump from Sonnet 5. Terminal-Bench climbs by 60 points, Chartography roughly quadruples, and GDPval-AA rises by nearly 400 Elo points.
Table 1 - Sonnet 5.5 benchmarks as published by Anthropic, as of September 2026. Anthropic notes that the GPT-6 Sol knowledge work and Chartography scores may not yet reflect an OpenAI image-understanding fix. Anthropic's Terminal-Bench and CursorBench charts compare against GPT-5.6 Sol, since no public GPT-6 Sol scores were available.
Each benchmark predicts a different kind of real work
A benchmark score only helps if you know what job it stands for. Several of these tests measure things that matter to a business app, while others matter mostly to engineering teams.
Table 2 - What each Sonnet 5.5 benchmark tests, based on Anthropic's descriptions, as of September 2026.
Two scales appear in Table 1. Percentages show how many tasks the model solved. Elo scores come from head-to-head comparisons. The 2-point GDPval-AA gap (1844 vs 1846) is small in practical terms, though the published figures don't include the error margins needed to call it a statistical tie.
Coding shows Sonnet 5.5's biggest jump over Sonnet 5
Coding is where Sonnet 5.5 improved most. Anthropic reports that at High effort on FrontierCode, it scores 10 points above Sonnet 5 at the same setting, for about one-fifteenth of the cost per task.
Early testers also noted that it batches tool calls together more often than Sonnet 5. That means fewer steps per job and a smaller bill.
1. Terminal-Bench 4.0 climbs from 10.3% to 70.6%
Terminal-Bench 4.0 is the one launch test where Sonnet 5.5 beats Opus 5.5. It scored 70.6% against Opus 5.5's 66.4%, and that Opus figure is its best run, at Xhigh effort.
The gain over Sonnet 5 is the largest on the page. According to Anthropic, Sonnet 5.5 at Medium effort, the default in the Claude apps, beats Sonnet 5's best Terminal-Bench score for less than a tenth of the cost per task.
2. FrontierCode drops at Max effort, and the reason matters
Sonnet 5.5 scores lower on FrontierCode at Max effort (46.2%) than at Xhigh (52.1%). More effort usually means a higher score, so the reversal is worth understanding.
Anthropic's footnote explains it. At Max, the model more often ran a code-review skill that splits work across many subagents. In the two cases Cognition examined, this caused a timeout or added edits beyond the task, which FrontierCode penalizes. At Xhigh, Sonnet 5.5 finishes 2.3 points behind Opus 5.5 and ahead of GPT-6 Sol's 49.3%. At Max, it falls below GPT-6 Sol.
3. CursorBench lands 2.3 points behind Opus 5.5
On CursorBench 4.0, Sonnet 5.5 scored 55.5% against Opus 5.5's 57.8%. That puts it second of the models Anthropic reported, and up from 34.1% for Sonnet 5.
CursorBench draws on real coding sessions, so it is the closest proxy here for day-to-day feature work. A 2.3-point gap is narrow, which suggests routine feature work will often land at similar quality on either model. Check it against your own tasks before switching.
We recommend our Opus 5.5 vs Sonnet 5 comparison if you're deciding which tier to run rather than just reading the scores.
Knowledge work and computer use scores sit almost level with Opus 5.5
Sonnet 5.5 nearly matches Opus 5.5 on business tasks. It scored 1844 on GDPval-AA against 1846 for Opus 5.5, and about 400 points above Sonnet 5. On AA-Briefcase, the gap is 11 Elo points (1811 vs 1822).
Computer use and chart reading follow the same pattern. OSWorld 2.1 sits at 80.1% against 81.8%, and Chartography at 61.6% against 64.4%. Chartography shows the largest relative gain in the table, since Sonnet 5 scored 15.6% on it.
Anthropic also reports a result from an internal test. It gave Sonnet 5.5 a public company's quarterly earnings materials and a slide template, and asked for a 10-slide operating review. Two experts judged the first draft ready to send. It is a single test, but it matches the knowledge work scores.
Independent testing tells a more mixed story than the launch post
Independent results confirm that Sonnet 5.5 is near the top, but they complicate the cost claims. Artificial Analysis runs every model through the same set of evaluations, which makes its figures the fairest way to compare models.
1. Artificial Analysis ranks Sonnet 5.5 second, 2 points behind Opus 5.5
Sonnet 5.5 at Max effort scores 56 on the Artificial Analysis Intelligence Index. That places it second, behind only Opus 5.5 at 58, and 18 points above Sonnet 5.
It also outscores Claude Fable 5.1, which sits at 53 on the same index. On AutomationBench-AA, Sonnet 5.5 edges Opus 5.5, 71% to 70%.
2. Terminal-Bench drops from 70.6% to 64% on a different harness
Artificial Analysis measured 64% on Terminal-Bench 4.0, compared with Anthropic's 70.6%. The two figures come from separate runs by different evaluators with different test setups (called harnesses), and neither source breaks down the 6.6-point gap. That is why vendor and independent figures should never share a table without a label.
The ranking holds on both. The independent run still puts Sonnet 5.5 slightly above Opus 5.5 and GPT-6 Astra, which both scored 60%.
3. Max effort uses more output tokens than any model Artificial Analysis has tested
At Max effort, Sonnet 5.5 used about 193k output tokens per Intelligence Index task. No model the firm has tested used more. It is roughly 60% more than Opus 5.5 or Sonnet 5 at Max, and about seven times GPT-6 Astra.
Customer reports from Anthropic's launch post point the other way, and both can be true. Balyasny Asset Management saw about 121k tokens per answer on 2,441 finance tasks, against 497k for Sonnet 5. Slack reported about 14% fewer output tokens on its Slackbot evaluations. These teams run production settings, not Max effort.
Table 3 - Vendor vs independent Sonnet 5.5 results, as of September 2026. Artificial Analysis tested a pre-release build and plans to rerun some evaluations.
One caveat applies to all the independent numbers. They come from a pre-release deployment that had a bug affecting structured outputs. The bug is fixed, Anthropic expects the effect to be small, and Artificial Analysis plans to rerun the affected evaluations.
Sonnet 5.5 still trails Opus 5.5 on factual knowledge
Factual recall is the clearest area where Sonnet 5.5 falls short of Opus 5.5. On AA-Omniscience, Artificial Analysis measured 54% factual accuracy for Sonnet 5.5 against 66% for Opus 5.5. That 12-point gap is far wider than any gap in Anthropic's launch table.
Artificial Analysis also measured a lower hallucination rate for Sonnet 5.5, at 47% against 59% for Opus 5.5. So it answers fewer factual questions correctly, yet shows a lower hallucination rate on this test. The same independent suite places it about six points behind Opus 5.5 on Humanity's Last Exam and SciCode.
For apps that answer questions from general knowledge, such as a research assistant or a domain lookup tool, this gap matters more than any coding score. Ground those answers in your own documents, or use Opus 5.5.
Token prices didn't change, but cost per task swings 18x with effort
Sonnet 5.5 costs the same per token as Sonnet 5, and half as much as Opus 5.5. What you actually pay depends on how many tokens each task uses, and the effort setting controls that.
1. Sonnet 5.5 pricing matches Sonnet 5 at half the cost of Opus 5.5
Table 4 - Claude Sonnet 5.5 vs Opus 5.5 API pricing, pricing as of September 2026.
The Claude Platform docs list a 1-hour cache write at $4 and a 50% Batch API discount. Sonnet 5.5 has a 1M-token context window and 128K max output, rising to 300K on the Batch API in beta. It takes text and images as input, and its reliable knowledge cutoff is June 2026. For the Opus side of the pricing, see our Opus 5.5 pricing breakdown.
2. Effort level sets both the score and the bill
Each step up in effort buys a higher score for a steeply higher price. On Artificial Analysis's release page, cost per task rises about 18 times from Low to Max, while the index score rises by 20 points.
Table 5 - Sonnet 5.5 score, cost, and speed by effort level, independently measured by Artificial Analysis, as of September 2026.
The steps are not evenly priced. Moving from High to Xhigh adds five points for about 2.5 times the cost. Moving from Xhigh to Max adds four points for nearly three times the cost again.
Defaults differ by surface. The Claude apps and Claude Code use Medium, and the Claude Platform API uses High. For cost-sensitive work, test Medium and High before moving up, since cost climbs sharply at the top two settings.
3. Anthropic's "30% cheaper" and the independent "50% pricier" are both true
The two cost claims measure different things. Anthropic's "up to 30% less per task" comes from its own testing of typical work. Artificial Analysis's figure of about 50% more than Sonnet 5 is measured at Max effort, where Sonnet 5.5 burns the most tokens.
Table 5 closes the gap. At Medium, Sonnet 5.5 scores 41 on the index, already above the 38 that Sonnet 5 reached at its best, for an estimated $0.59 per task. Only at the top settings does its cost pass Sonnet 5's. These are benchmark workload figures, so treat them as a guide to direction rather than a forecast of your own bill.
We recommend our GPT-6 Sol benchmarks breakdown for the closest comparison at this tier.
Sonnet 5.5 is the better pick for well-scoped, high-volume work
Sonnet 5.5 wins wherever the task is clear and the volume is high. Anthropic's own guidance matches the data: Opus 5.5 stays clearly stronger at complex, open-ended work that needs sustained judgment.
Table 6 - When to choose Sonnet 5.5 or Opus 5.5, based on published benchmarks and tester reports, as of September 2026.
One behavior change is worth knowing before you switch. Sonnet 5.5 is the first Sonnet model to ship with cyber safeguards. Higher-risk cybersecurity requests visibly fall back to Sonnet 5, and routine bug fixing is unaffected. In independent testing, the fallback triggered in about 0.1% of tasks.
If your work leans toward judgment-heavy tasks, start with Opus 5.5 and where its extra cost pays off.
Pick Sonnet 5.5 for everyday work and save Opus 5.5 for judgment calls
The Sonnet 5.5 benchmarks support a clear split. Sonnet 5.5 matches Opus 5.5 closely on coding, knowledge work, and computer use, and leads it on agentic terminal tasks, at half the token price. It trails on factual recall and open-ended judgment. Start at Medium or High effort, and move higher only when your own tasks show the gain is worth the cost.
Sonnet 5.5 is available on Emergent through the Universal LLM Key. The apps you build can call it without a separate Anthropic account or API key, with usage billed through Emergent Credits. It fits the high-volume, well-scoped jobs where it wins on price, like a client portal that drafts support replies or an internal tool that turns spreadsheets into reports. When you create a custom agent, you choose the language model it reasons with at setup.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







