HomeLearn

Sonnet 5.5 vs Sonnet 5: What Changed and Should You Upgrade?

Sonnet 5.5 vs Sonnet 5: same token price, an 18-point Intelligence Index jump, and lower benchmark cost per task at 4 of 5 effort levels. Should you upgrade?

Divit Bhat
Written by
Divit Bhat
Bhavyadeep
Reviewed by
Bhavyadeep
Last updated: 
September 29, 2026
0
 min read
Select Emergent as your Preferred news source
Table of Contents

TL;DR

  • Sonnet 5.5 vs Sonnet 5 is close to a straight upgrade: both cost $2 per million input tokens and $10 per million output tokens, and Sonnet 5.5 scores higher on every benchmark in Anthropic's launch comparison.
  • Artificial Analysis scores Sonnet 5.5 at 56 on its Intelligence Index, 18 points above Sonnet 5's best of 38.
  • On Artificial Analysis's benchmark workload, Sonnet 5.5 costs less per task than Sonnet 5 at four of five effort levels. At Low, it scores 36 for $0.41 per task, close to Sonnet 5's best score at about a twelfth of the $5.09 cost.
  • Max effort is the exception: Sonnet 5.5 costs about 49% more per task there, while scoring 18 points higher.
  • Sonnet 5 stays available until at least June 30, 2027, so you can switch on your own schedule. Re-test your effort settings when you do.

‍

Anthropic kept the price of Sonnet 5.5 identical to Sonnet 5 and raised nearly every score. For most workloads, that makes the upgrade decision simple.

The details are in how the two models spend tokens. Anthropic's testing and independent benchmarks both show it using fewer tokens per task at the settings most teams run, so it usually costs less per task. Push it to Max effort, though, and it becomes the more expensive model.

Beyond that, the choice comes down to one caveat and a short list of behavior changes. For a full breakdown of each score, see the Sonnet 5.5 benchmarks.

Sonnet 5.5 vs Sonnet 5 at a glance: same price, much higher scores

The two models share a price, a tokenizer, and a 1M-token context window. The same text produces the same token counts on both, so bills compare directly.

Spec Claude Sonnet 5.5 Claude Sonnet 5
Input / output price per 1M tokens $2 / $10 $2 / $10
Caching and batch pricing Same as Sonnet 5 Same as Sonnet 5.5
Tokenizer Same as Sonnet 5 Same as Sonnet 5.5
Context window 1M tokens 1M tokens
Minimum prompt length for caching 512 tokens 1,024 tokens
Max output 128K tokens Not listed in current docs
Reliable knowledge cutoff June 2026 Not listed in current docs
Released September 28, 2026 June 2026
Retirement Not sooner than September 28, 2027 Not sooner than June 30, 2027

Table 1 - Claude Sonnet 5.5 vs Sonnet 5 specs from Anthropic's docs and Artificial Analysis, as of September 2026.

Anthropic's "What's new" docs confirm the matching prices, tokenizer, and caching minimum. Sonnet 5 no longer appears in Anthropic's current models table, which is why two of its rows are blank.

Sonnet 5.5 improves on every launch benchmark, most on coding and charts

Sonnet 5.5 beats Sonnet 5 on all eight benchmarks in Anthropic's launch post. On the percentage-scored tests, the gains range from about 10 points to more than 60.

Benchmark Sonnet 5.5 Sonnet 5 Gain
Terminal-Bench 4.0 (agentic coding) 70.6% 10.3% +60.3 points
Chartography (chart recognition, no tools) 61.6% 15.6% +46.0 points
OSWorld 2.1 (computer use, partial credit) 80.1% 57.0% +23.1 points
CursorBench 4.0 (agentic coding) 55.5% 34.1% +21.4 points
FrontierCode 1.1 Main (agentic coding) 52.1% (Xhigh), 46.2% (Max) 42.4% +9.7 points at Xhigh
Humanity's Last Exam (with tools) 64.5% 54.9% +9.6 points
AA-Briefcase v1.1 (knowledge work, Elo) 1811 1359 +452 Elo
GDPval-AA v2.1 (knowledge work, Elo) 1844 1449 +395 Elo

Table 2 - Sonnet 5.5 vs Sonnet 5 benchmarks as published by Anthropic, as of September 2026. Artificial Analysis ran GDPval-AA and AA-Briefcase.

The largest jumps land where Sonnet 5 was weakest. Terminal-Bench measures long, multi-step work in a command-line setting, and on Anthropic's setup, Sonnet 5 solved about one task in ten. The new model solves seven in ten.

Chartography nearly quadruples, which matters for any app that reads dashboards, reports, or screenshots. Knowledge work scores rose by roughly 400 to 450 Elo points, a large margin on a head-to-head scale.

Independent testing shows an 18-point jump on the Intelligence Index

Artificial Analysis confirms the gain on its own standardized tests. At Max effort, Sonnet 5.5 scores 56 on its Intelligence Index, against 38 for Sonnet 5 at Max.

That moved Sonnet 5.5 to second place on the index at launch, behind only Claude Opus 5.5. On the firm's own Terminal-Bench 4.0 run, Sonnet 5.5 scored 64%, about 50 points above Sonnet 5.

These results come with a caveat. The firm tested a pre-release Sonnet 5.5 deployment that had a structured-output bug, since fixed, and plans to rerun the affected evaluations.

Same token price, lower benchmark cost per task at every effort level except Max

On Artificial Analysis's weighted benchmark workload, Sonnet 5.5 is cheaper per task than Sonnet 5 at Low, Medium, High, and Xhigh effort. It also scores 12 to 18 points higher at each of those levels. The table pairs both models' results from the Artificial Analysis release pages.

Effort Sonnet 5.5 score / cost per task Sonnet 5 score / cost per task Sonnet 5.5 change
Low 36 / $0.41 24 / $0.51 +12 points, 20% cheaper
Medium 41 / $0.59 28 / $1.00 +13 points, 41% cheaper
High 47 / $1.08 32 / $1.79 +15 points, 40% cheaper
Xhigh 52 / $2.74 34 / $2.87 +18 points, 5% cheaper
Max 56 / $7.60 38 / $5.09 +18 points, 49% more expensive

Table 3 - Sonnet 5.5 vs Sonnet 5 Intelligence Index score and estimated cost per weighted index task by effort level, independently measured by Artificial Analysis, as of September 2026. Figures reflect a benchmark workload, not a production bill.

1. Sonnet 5.5 at Low nearly matches Sonnet 5 at its best for a twelfth of the cost

The most striking row pair is across effort levels. Sonnet 5.5 at Low scores 36 for $0.41 per task, two points short of Sonnet 5's best score of 38 at Max, which costs $5.09.

At Medium, Sonnet 5.5 passes Sonnet 5's best outright, scoring 41 for $0.59. Anthropic's own testing points the same way. It reports that at Medium effort, Sonnet 5.5 beats Sonnet 5's best Terminal-Bench score for less than a tenth of the cost per task.

2. Max effort costs 49% more than Sonnet 5 at Max

Max is the one setting where Sonnet 5.5 costs more. It spends $7.60 per index task against $5.09 for Sonnet 5, because it writes far more output at that level.

At Max, Sonnet 5.5 used about 193k output tokens per index task, the most Artificial Analysis had recorded for any model. That is roughly 60% more than Sonnet 5 used at Max. The extra spend buys 18 more index points, but few workloads need it.

3. Testers report far fewer tokens per job at everyday settings

Customer-reported results in Anthropic's launch post show the other side of the token story. These are private evaluations, not public benchmarks. Balyasny Asset Management saw about 121k tokens per answer across 2,441 finance tasks, against 497k for Sonnet 5.

Slack reported about 14% fewer output tokens on its Slackbot evaluations, and Box reported 2.4 times faster runs with 12% fewer total tokens. These teams reported results from their own workloads rather than Max-effort benchmark runs. Anthropic sums it up as up to 30% lower cost per task than Sonnet 5 for most work.

Suggested read: Claude Sonnet 5.5 Pricing: API Rates, Plans, and What a Task Really Costs

Sonnet 5.5 is faster at every effort level

Anthropic says Sonnet 5.5 generates output more than 30% faster than Sonnet 5. The independent speed data shows a larger gap.

Across effort levels, Artificial Analysis measured Sonnet 5.5 at 85 to 139 output tokens per second, against 55 to 75 for Sonnet 5. The gap runs from 43% faster at High to 86% faster at Medium.

Speed per token still isn't time per task. At Max, Sonnet 5.5's heavier output can offset its faster stream. At Low through High, it is both faster per token and lighter on tokens, so jobs should finish sooner.

Four changes to plan for before you switch

Sonnet 5.5 behaves differently from Sonnet 5 in ways that show up without any code change. Anthropic's docs list these four as the ones most likely to affect an existing app.

1. Effort levels are recalibrated

An effort setting on Sonnet 5.5 doesn't produce the same amount of thinking it did on Sonnet 5. Anthropic advises re-testing rather than carrying a setting over. Its starting points are High for most work and Medium for well-specified agent tasks or latency-sensitive chat.

2. Progress notes between tool calls come back differently

When Sonnet 5.5 pauses between tool calls, longer progress notes come back inside its thinking blocks instead of as plain text. An app that shows those notes to users can go quiet between steps until it is configured to display them. Nothing errors, so this is easy to miss in testing.

3. Some requests fall back to Sonnet 5

Sonnet 5.5 is the first Sonnet model to launch with cyber safeguards. When the API's fallback option is turned on, requests that the safeguards decline as higher-risk cybersecurity work can be retried on Sonnet 5. Routine bug fixing is unaffected. In independent testing, the fallback triggered in about 0.1% of tasks.

4. Shorter prompts can now be cached

The minimum prompt length for caching drops from 1,024 tokens on Sonnet 5 to 512 on Sonnet 5.5. Apps with short, repeated instructions can now cache them. Cached input tokens cost 90% less when reused, while output tokens and cache writes are billed as usual.

Conversation history also carries forward. Sonnet 5.5 can read Sonnet 5's thinking blocks, so a conversation that moves up from Sonnet 5 keeps its earlier reasoning. For teams with custom code, Anthropic also lists five breaking API changes in its migration guide. They cover how up-front thinking is turned off, the removal of forced tool use, rules against editing earlier conversation history, an updated computer-use tool, and which models can act as an advisor.

Sonnet 5 stays available until at least June 2027

There is no deadline forcing the switch. Anthropic's model deprecations page lists Sonnet 5 as active, with retirement not sooner than June 30, 2027.

That leaves time to test before moving production work. A few situations justify staying on Sonnet 5 a little longer:

  • Your integration depends on a feature Sonnet 5.5 dropped, such as forced tool use, and the code hasn't been updated yet.
  • Your prompts and effort settings are heavily tuned, and you need time to re-test them on the recalibrated effort levels.
  • Your workload runs at Max effort under a fixed per-task budget. In that case, test Sonnet 5.5 at Xhigh first, where it scores 52 for $2.74 against Sonnet 5's 38 for $5.09.

Upgrade to Sonnet 5.5 for almost every workload

Sonnet 5.5 is the better model at the same price, and it costs less per task at the settings most teams use. The decision table covers the common cases.

Situation Recommendation Why
Running Sonnet 5 at Low, Medium, or High Upgrade now Higher scores and 20% to 41% lower cost per task
Running Sonnet 5 at Xhigh Upgrade now 18 points higher for 5% less per task
Running Sonnet 5 at Max Upgrade and drop to Xhigh Sonnet 5.5 at Xhigh beats Sonnet 5 at Max by 14 points for about half the cost
Apps that read charts, dashboards, or screenshots Upgrade now Chartography rose from 15.6% to 61.6%
Long multi-step agent tasks Upgrade now Terminal-Bench rose from 10.3% to 70.6%
Custom integrations touching any of the five breaking changes Upgrade after code changes Forced tool use, disabled thinking, edited history, and the older computer-use tool can return errors on Sonnet 5.5

Table 4 - When to move from Sonnet 5 to Sonnet 5.5, based on Anthropic's published data and Artificial Analysis results, as of September 2026.

If Sonnet 5.5 still falls short at practical settings, compare Opus 5.5 with Sonnet 5.5 at Max on your own tasks before committing to either. The Sonnet 5.5 vs Opus 5.5 comparison covers where that line falls.

Suggested read: Claude Opus 5.5 vs Claude Sonnet 5: The Ultimate Comparison

Get Sonnet 5.5's gains without touching an API key

Sonnet 5.5 is a strong default upgrade from Sonnet 5, though not a blind drop-in replacement. It keeps the same price, with much higher scores, faster output, and lower benchmark cost per task at every effort level below Max. Switch on your own schedule before Sonnet 5 retires, re-test your effort settings, and skip Max unless the work truly needs it.

Sonnet 5.5 is now live on Emergent. To add it to the apps you build, you can use your own Anthropic API key or the Universal LLM Key, which skips the separate account and key setup.

Usage is billed through Emergent Credits. When you create a custom agent, you choose the language model it reasons with at setup. That makes Sonnet 5.5 a practical fit for a client portal that drafts support replies or an internal tool that turns spreadsheets into reports.

Start Building on Emergent.

Was this article helpful?
About the writer

Divit Bhat is a product and growth writer at Emergent, specializing in AI-powered app building, no code platforms, and modern software workflows. He creates practical guides and tutorials to help founders, enterprises and teams build, automate, and scale products with AI.

Cta image

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free
Share this article:

Frequently Asked Questions

Your Questions, Answered

Is Sonnet 5.5 better than Sonnet 5?
Yes. Sonnet 5.5 scores higher on all eight benchmarks in Anthropic's launch comparison, including a jump from 10.3% to 70.6% on Terminal-Bench 4.0. On the Artificial Analysis Intelligence Index, it scores 56 against 38 for Sonnet 5. It also runs faster and costs less per task at most effort levels.
Is Sonnet 5.5 more expensive than Sonnet 5?
No, the token prices are identical at $2 per million input tokens and $10 per million output tokens, including caching and batch rates. Per task, Sonnet 5.5 is cheaper at Low through Xhigh effort on Artificial Analysis's benchmark workload. At Max, Artificial Analysis measured $7.60 per task against $5.09 for Sonnet 5, about 49% more.
Should I upgrade from Sonnet 5 to Sonnet 5.5?
For almost every workload, yes. It costs the same per token, and on benchmark tasks it delivers higher scores at lower cost per task below Max effort. Re-test your effort settings after switching, since Anthropic recalibrated them. Custom integrations affected by any of the five breaking API changes, such as forced tool use or disabled thinking, need code changes first.
When will Sonnet 5 be retired?
Anthropic's model deprecations page lists Sonnet 5 as active, with retirement not sooner than June 30, 2027. That gives teams time to test Sonnet 5.5 before moving production work. The newer model is listed with retirement not sooner than September 28, 2027.
Is Sonnet 5.5 faster than Sonnet 5?
Yes. Anthropic says Sonnet 5.5 generates output more than 30% faster. Artificial Analysis measured 85 to 139 output tokens per second for Sonnet 5.5 against 55 to 75 for Sonnet 5 across effort levels. At Max effort, Sonnet 5.5's heavier output can offset part of that speed advantage.
Do my Sonnet 5 effort settings carry over to Sonnet 5.5?
Not reliably. Anthropic recalibrated the effort levels, so the same setting produces a different amount of thinking on Sonnet 5.5. It recommends re-testing rather than copying settings over, starting at High for most work and Medium for well-specified agent tasks or chat.
Start Building
on Emergent today
Try Emergent

https://api.linear.app/graphql