HomeLearn

Claude Sonnet 5.5 Benchmarks: Scores, Real Costs, and Where It Trails Opus 5.5

Claude Sonnet 5.5 benchmarks explained: 70.6% on Terminal-Bench 4.0, #2 on Artificial Analysis, and why cost per task rises 18x from Low to Max effort.

Bhavyadeep
Written by
Bhavyadeep
Priyanka Singh
Reviewed by
Priyanka Singh
Last updated: 
September 29, 2026
0
 min read
Select Emergent as your Preferred news source
Table of Contents

TL;DR

  • Sonnet 5.5 benchmarks trail Claude Opus 5.5 on seven of eight launch tests, by 3.2 percentage points or less on the percentage-scored ones, and lead on one: Terminal-Bench 4.0 (70.6% vs 66.4%).
  • Artificial Analysis independently ranks Sonnet 5.5 second on its Intelligence Index at 56, two points behind Opus 5.5.
  • Token prices are unchanged from Sonnet 5 at $2 input and $10 output per million tokens, half of Opus 5.5.
  • Cost per task depends on effort: on Artificial Analysis's weighted Intelligence Index workload, it rises from $0.41 at Low to $7.60 at Max.
  • Run Sonnet 5.5 at Medium or High for well-scoped, high-volume work. Keep Opus 5.5 for open-ended work that needs judgment and deep factual recall.

‍

Claude Sonnet 5.5 is the closest a Sonnet model has come to Opus quality. Anthropic's launch table has it within 3.2 percentage points of Opus 5.5 on every percentage-scored test it trails, and ahead on agentic terminal coding. The catch is in the cost column.

If you are choosing a model for an app, a support workflow, or an internal tool, the headline Sonnet 5.5 benchmarks tell half the story. The effort setting you pick can move the price of one task from under $0.50 to over $7.

Independent testing also disagrees with the launch post on cost, and that gap decides whether Sonnet 5.5 beats the Opus 5.5 benchmarks on value per dollar.

Sonnet 5.5 benchmarks trail Opus 5.5 on 7 of 8 tests, but never by much

Sonnet 5.5 scores close to Opus 5.5 across coding, knowledge work, reasoning, computer use, and chart reading. Anthropic's launch post shades the top score in each row, and Sonnet 5.5 holds that spot only on Terminal-Bench 4.0.

The bigger story is the jump from Sonnet 5. Terminal-Bench climbs by 60 points, Chartography roughly quadruples, and GDPval-AA rises by nearly 400 Elo points.

Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Terminal-Bench 4.0 (agentic coding) 70.6% 10.3% 66.4% (Xhigh) Not reported
FrontierCode 1.1 Main (agentic coding) 52.1% (Xhigh), 46.2% (Max) 42.4% 54.4% 49.3%
CursorBench 4.0 (agentic coding) 55.5% 34.1% 57.8% Not reported
GDPval-AA v2.1 (knowledge work, Elo) 1844 1449 1846 1487
AA-Briefcase v1.1 (knowledge work, Elo) 1811 1359 1822 1483
Humanity's Last Exam (with tools) 64.5% 54.9% 67.7% Not reported
OSWorld 2.1 (computer use, partial credit) 80.1% 57.0% 81.8% Not reported
Chartography (chart recognition, no tools) 61.6% 15.6% 64.4% 53.6%

Table 1 - Sonnet 5.5 benchmarks as published by Anthropic, as of September 2026. Anthropic notes that the GPT-6 Sol knowledge work and Chartography scores may not yet reflect an OpenAI image-understanding fix. Anthropic's Terminal-Bench and CursorBench charts compare against GPT-5.6 Sol, since no public GPT-6 Sol scores were available.

Each benchmark predicts a different kind of real work

A benchmark score only helps if you know what job it stands for. Several of these tests measure things that matter to a business app, while others matter mostly to engineering teams.

Benchmark What it measures What it predicts for your work
Terminal-Bench 4.0 Complex, multi-step professional tasks in a command-line interface Whether the model can finish a long technical job without step-by-step hand-holding
FrontierCode 1.1 Whether a code change could be merged with no human edits; out-of-scope changes are penalized Whether the model stays inside the task you gave it
CursorBench 4.0 Tasks drawn from real Cursor coding sessions Everyday feature work and bug fixes
GDPval-AA v2.1 Real-world tasks across 44 occupations and nine industries, scored as Elo Quality of business deliverables like reports, analyses, and plans
AA-Briefcase v1.1 Long-horizon knowledge work, scored as Elo Multi-step office work that runs across many files and steps
Humanity's Last Exam Multidisciplinary expert-level reasoning, with tools allowed Hard analytical questions at the edge of expert knowledge
OSWorld 2.1 Operating software through a computer interface Agents that click through apps and forms on your behalf
Chartography Reading and interpreting charts from images, with no tools Dashboards, reports, and screenshots as inputs

Table 2 - What each Sonnet 5.5 benchmark tests, based on Anthropic's descriptions, as of September 2026.

Two scales appear in Table 1. Percentages show how many tasks the model solved. Elo scores come from head-to-head comparisons. The 2-point GDPval-AA gap (1844 vs 1846) is small in practical terms, though the published figures don't include the error margins needed to call it a statistical tie.

Coding shows Sonnet 5.5's biggest jump over Sonnet 5

Coding is where Sonnet 5.5 improved most. Anthropic reports that at High effort on FrontierCode, it scores 10 points above Sonnet 5 at the same setting, for about one-fifteenth of the cost per task.

Early testers also noted that it batches tool calls together more often than Sonnet 5. That means fewer steps per job and a smaller bill.

1. Terminal-Bench 4.0 climbs from 10.3% to 70.6%

Terminal-Bench 4.0 is the one launch test where Sonnet 5.5 beats Opus 5.5. It scored 70.6% against Opus 5.5's 66.4%, and that Opus figure is its best run, at Xhigh effort.

The gain over Sonnet 5 is the largest on the page. According to Anthropic, Sonnet 5.5 at Medium effort, the default in the Claude apps, beats Sonnet 5's best Terminal-Bench score for less than a tenth of the cost per task.

2. FrontierCode drops at Max effort, and the reason matters

Sonnet 5.5 scores lower on FrontierCode at Max effort (46.2%) than at Xhigh (52.1%). More effort usually means a higher score, so the reversal is worth understanding.

Anthropic's footnote explains it. At Max, the model more often ran a code-review skill that splits work across many subagents. In the two cases Cognition examined, this caused a timeout or added edits beyond the task, which FrontierCode penalizes. At Xhigh, Sonnet 5.5 finishes 2.3 points behind Opus 5.5 and ahead of GPT-6 Sol's 49.3%. At Max, it falls below GPT-6 Sol.

3. CursorBench lands 2.3 points behind Opus 5.5

On CursorBench 4.0, Sonnet 5.5 scored 55.5% against Opus 5.5's 57.8%. That puts it second of the models Anthropic reported, and up from 34.1% for Sonnet 5.

CursorBench draws on real coding sessions, so it is the closest proxy here for day-to-day feature work. A 2.3-point gap is narrow, which suggests routine feature work will often land at similar quality on either model. Check it against your own tasks before switching.

We recommend our Opus 5.5 vs Sonnet 5 comparison if you're deciding which tier to run rather than just reading the scores.

Knowledge work and computer use scores sit almost level with Opus 5.5

Sonnet 5.5 nearly matches Opus 5.5 on business tasks. It scored 1844 on GDPval-AA against 1846 for Opus 5.5, and about 400 points above Sonnet 5. On AA-Briefcase, the gap is 11 Elo points (1811 vs 1822).

Computer use and chart reading follow the same pattern. OSWorld 2.1 sits at 80.1% against 81.8%, and Chartography at 61.6% against 64.4%. Chartography shows the largest relative gain in the table, since Sonnet 5 scored 15.6% on it.

Anthropic also reports a result from an internal test. It gave Sonnet 5.5 a public company's quarterly earnings materials and a slide template, and asked for a 10-slide operating review. Two experts judged the first draft ready to send. It is a single test, but it matches the knowledge work scores.

Independent testing tells a more mixed story than the launch post

Independent results confirm that Sonnet 5.5 is near the top, but they complicate the cost claims. Artificial Analysis runs every model through the same set of evaluations, which makes its figures the fairest way to compare models.

1. Artificial Analysis ranks Sonnet 5.5 second, 2 points behind Opus 5.5

Sonnet 5.5 at Max effort scores 56 on the Artificial Analysis Intelligence Index. That places it second, behind only Opus 5.5 at 58, and 18 points above Sonnet 5.

It also outscores Claude Fable 5.1, which sits at 53 on the same index. On AutomationBench-AA, Sonnet 5.5 edges Opus 5.5, 71% to 70%.

2. Terminal-Bench drops from 70.6% to 64% on a different harness

Artificial Analysis measured 64% on Terminal-Bench 4.0, compared with Anthropic's 70.6%. The two figures come from separate runs by different evaluators with different test setups (called harnesses), and neither source breaks down the 6.6-point gap. That is why vendor and independent figures should never share a table without a label.

The ranking holds on both. The independent run still puts Sonnet 5.5 slightly above Opus 5.5 and GPT-6 Astra, which both scored 60%.

3. Max effort uses more output tokens than any model Artificial Analysis has tested

At Max effort, Sonnet 5.5 used about 193k output tokens per Intelligence Index task. No model the firm has tested used more. It is roughly 60% more than Opus 5.5 or Sonnet 5 at Max, and about seven times GPT-6 Astra.

Customer reports from Anthropic's launch post point the other way, and both can be true. Balyasny Asset Management saw about 121k tokens per answer on 2,441 finance tasks, against 497k for Sonnet 5. Slack reported about 14% fewer output tokens on its Slackbot evaluations. These teams run production settings, not Max effort.

Claim Anthropic (vendor-reported) Artificial Analysis (independent) Why they differ
Terminal-Bench 4.0 70.6% 64% Separate runs by different evaluators and setups
Lead over Opus 5.5 on Terminal-Bench 4.2 points About 4 points Both agree Sonnet 5.5 leads
Overall standing Comparable to Opus 5.5 at Max on several tests #2 on the Intelligence Index, 56 vs 58 Consistent
Cost per task vs Sonnet 5 Up to 30% less About 50% more at Max Anthropic's claim covers its tested workloads; AA's figure is its weighted index workload at Max
Token use Far fewer tokens for the same work Highest token use measured, at Max Typical settings vs Max effort

Table 3 - Vendor vs independent Sonnet 5.5 results, as of September 2026. Artificial Analysis tested a pre-release build and plans to rerun some evaluations.

One caveat applies to all the independent numbers. They come from a pre-release deployment that had a bug affecting structured outputs. The bug is fixed, Anthropic expects the effect to be small, and Artificial Analysis plans to rerun the affected evaluations.

Sonnet 5.5 still trails Opus 5.5 on factual knowledge

Factual recall is the clearest area where Sonnet 5.5 falls short of Opus 5.5. On AA-Omniscience, Artificial Analysis measured 54% factual accuracy for Sonnet 5.5 against 66% for Opus 5.5. That 12-point gap is far wider than any gap in Anthropic's launch table.

Artificial Analysis also measured a lower hallucination rate for Sonnet 5.5, at 47% against 59% for Opus 5.5. So it answers fewer factual questions correctly, yet shows a lower hallucination rate on this test. The same independent suite places it about six points behind Opus 5.5 on Humanity's Last Exam and SciCode.

For apps that answer questions from general knowledge, such as a research assistant or a domain lookup tool, this gap matters more than any coding score. Ground those answers in your own documents, or use Opus 5.5.

Token prices didn't change, but cost per task swings 18x with effort

Sonnet 5.5 costs the same per token as Sonnet 5, and half as much as Opus 5.5. What you actually pay depends on how many tokens each task uses, and the effort setting controls that.

1. Sonnet 5.5 pricing matches Sonnet 5 at half the cost of Opus 5.5

Price per 1M tokens Claude Sonnet 5.5 Claude Opus 5.5
Input $2 $4
Output $10 $20
Cache reads $0.20 $0.20
Cache writes (5 minutes) $2.50 $5

Table 4 - Claude Sonnet 5.5 vs Opus 5.5 API pricing, pricing as of September 2026.

The Claude Platform docs list a 1-hour cache write at $4 and a 50% Batch API discount. Sonnet 5.5 has a 1M-token context window and 128K max output, rising to 300K on the Batch API in beta. It takes text and images as input, and its reliable knowledge cutoff is June 2026. For the Opus side of the pricing, see our Opus 5.5 pricing breakdown.

2. Effort level sets both the score and the bill

Each step up in effort buys a higher score for a steeply higher price. On Artificial Analysis's release page, cost per task rises about 18 times from Low to Max, while the index score rises by 20 points.

Effort Intelligence Index Cost per index task Output speed
Low 36 $0.41 85 tokens/s
Medium 41 $0.59 104 tokens/s
High 47 $1.08 93 tokens/s
Xhigh 52 $2.74 113 tokens/s
Max 56 $7.60 139 tokens/s

Table 5 - Sonnet 5.5 score, cost, and speed by effort level, independently measured by Artificial Analysis, as of September 2026.

The steps are not evenly priced. Moving from High to Xhigh adds five points for about 2.5 times the cost. Moving from Xhigh to Max adds four points for nearly three times the cost again.

Defaults differ by surface. The Claude apps and Claude Code use Medium, and the Claude Platform API uses High. For cost-sensitive work, test Medium and High before moving up, since cost climbs sharply at the top two settings.

3. Anthropic's "30% cheaper" and the independent "50% pricier" are both true

The two cost claims measure different things. Anthropic's "up to 30% less per task" comes from its own testing of typical work. Artificial Analysis's figure of about 50% more than Sonnet 5 is measured at Max effort, where Sonnet 5.5 burns the most tokens.

Table 5 closes the gap. At Medium, Sonnet 5.5 scores 41 on the index, already above the 38 that Sonnet 5 reached at its best, for an estimated $0.59 per task. Only at the top settings does its cost pass Sonnet 5's. These are benchmark workload figures, so treat them as a guide to direction rather than a forecast of your own bill.

We recommend our GPT-6 Sol benchmarks breakdown for the closest comparison at this tier.

Sonnet 5.5 is the better pick for well-scoped, high-volume work

Sonnet 5.5 wins wherever the task is clear and the volume is high. Anthropic's own guidance matches the data: Opus 5.5 stays clearly stronger at complex, open-ended work that needs sustained judgment.

Use case Better pick Evidence
Support replies and ticket triage Sonnet 5.5 at Medium Zendesk reported tickets processed 20% faster
Bug fixes and scoped feature changes Sonnet 5.5 at High or Xhigh 2.3 points behind Opus 5.5 on CursorBench and FrontierCode
Documents, slides, and spreadsheets Sonnet 5.5 Within 2 Elo points of Opus 5.5 on GDPval-AA
Long multi-step agent tasks Sonnet 5.5 at Xhigh Leads Opus 5.5 on Terminal-Bench 4.0 in both vendor and independent runs
Open-ended planning and architecture Opus 5.5 Anthropic says Opus 5.5 is clearly stronger at sustained judgment
Fact-heavy answers from general knowledge Opus 5.5 12-point lead on AA-Omniscience factual accuracy
Anything you would run at Max effort Compare with Opus 5.5 first Anthropic says Sonnet 5.5 costs about the same as Opus 5.5 at high settings

Table 6 - When to choose Sonnet 5.5 or Opus 5.5, based on published benchmarks and tester reports, as of September 2026.

One behavior change is worth knowing before you switch. Sonnet 5.5 is the first Sonnet model to ship with cyber safeguards. Higher-risk cybersecurity requests visibly fall back to Sonnet 5, and routine bug fixing is unaffected. In independent testing, the fallback triggered in about 0.1% of tasks.

If your work leans toward judgment-heavy tasks, start with Opus 5.5 and where its extra cost pays off.

Pick Sonnet 5.5 for everyday work and save Opus 5.5 for judgment calls

The Sonnet 5.5 benchmarks support a clear split. Sonnet 5.5 matches Opus 5.5 closely on coding, knowledge work, and computer use, and leads it on agentic terminal tasks, at half the token price. It trails on factual recall and open-ended judgment. Start at Medium or High effort, and move higher only when your own tasks show the gain is worth the cost.

Sonnet 5.5 is available on Emergent through the Universal LLM Key. The apps you build can call it without a separate Anthropic account or API key, with usage billed through Emergent Credits. It fits the high-volume, well-scoped jobs where it wins on price, like a client portal that drafts support replies or an internal tool that turns spreadsheets into reports. When you create a custom agent, you choose the language model it reasons with at setup.

Start Building on Emergent.

Was this article helpful?
About the writer

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Cta image

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free
Share this article:

Frequently Asked Questions

Your Questions, Answered

What is Claude Sonnet 5.5's Terminal-Bench 4.0 score?
Anthropic reports 70.6% on Terminal-Bench 4.0, up from 10.3% for Sonnet 5 and ahead of Opus 5.5's 66.4% at Xhigh effort. Artificial Analysis measured 64% on its own harness, still slightly above the 60% it recorded for Opus 5.5 and GPT-6 Astra. Both sources place Sonnet 5.5 at or near the top.
Is Sonnet 5.5 better than Opus 5.5?
No, but it is close. Sonnet 5.5 trails Opus 5.5 on seven of eight launch benchmarks, by 3.2 percentage points or less on the percentage-scored ones, and leads on Terminal-Bench 4.0. Artificial Analysis ranks it 56 against Opus 5.5's 58. Opus 5.5 remains stronger on factual knowledge and open-ended work that needs sustained judgment.
How much does Claude Sonnet 5.5 cost?
Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Cache reads cost $0.20 and five-minute cache writes cost $2.50 per million tokens. Opus 5.5 costs twice as much at $4 input and $20 output. Batch API requests get a 50% discount.
Is Sonnet 5.5 cheaper per task than Sonnet 5?
At typical settings, yes. Anthropic reports up to 30% lower cost per task, and on Artificial Analysis's Intelligence Index, Sonnet 5.5 at Medium outscores Sonnet 5's best for an estimated $0.59 per task. At Max effort, it costs about 50% more per task than Sonnet 5, because it uses the most output tokens Artificial Analysis has measured.
Why does Sonnet 5.5 score lower at Max effort on FrontierCode?
At Max effort, Sonnet 5.5 more often ran a code-review skill that splits work across many subagents. In the cases Cognition examined, this caused timeouts or extra edits beyond the task, and FrontierCode penalizes out-of-scope changes. It scored 46.2% at Max against 52.1% at Xhigh.
What is Sonnet 5.5's context window?
Claude Sonnet 5.5 has a 1M-token context window and a 128K max output, rising to 300K output tokens on the Batch API in beta. It accepts text and image input and returns text. Its reliable knowledge cutoff is June 2026, and Anthropic announced it on September 28, 2026.
Start Building
on Emergent today
Try Emergent

https://api.linear.app/graphql