Anthropic kept the price of Sonnet 5.5 identical to Sonnet 5 and raised nearly every score. For most workloads, that makes the upgrade decision simple.
The details are in how the two models spend tokens. Anthropic's testing and independent benchmarks both show it using fewer tokens per task at the settings most teams run, so it usually costs less per task. Push it to Max effort, though, and it becomes the more expensive model.
Beyond that, the choice comes down to one caveat and a short list of behavior changes. For a full breakdown of each score, see the Sonnet 5.5 benchmarks.
Sonnet 5.5 vs Sonnet 5 at a glance: same price, much higher scores
The two models share a price, a tokenizer, and a 1M-token context window. The same text produces the same token counts on both, so bills compare directly.
Table 1 - Claude Sonnet 5.5 vs Sonnet 5 specs from Anthropic's docs and Artificial Analysis, as of September 2026.
Anthropic's "What's new" docs confirm the matching prices, tokenizer, and caching minimum. Sonnet 5 no longer appears in Anthropic's current models table, which is why two of its rows are blank.
Sonnet 5.5 improves on every launch benchmark, most on coding and charts
Sonnet 5.5 beats Sonnet 5 on all eight benchmarks in Anthropic's launch post. On the percentage-scored tests, the gains range from about 10 points to more than 60.
Table 2 - Sonnet 5.5 vs Sonnet 5 benchmarks as published by Anthropic, as of September 2026. Artificial Analysis ran GDPval-AA and AA-Briefcase.
The largest jumps land where Sonnet 5 was weakest. Terminal-Bench measures long, multi-step work in a command-line setting, and on Anthropic's setup, Sonnet 5 solved about one task in ten. The new model solves seven in ten.
Chartography nearly quadruples, which matters for any app that reads dashboards, reports, or screenshots. Knowledge work scores rose by roughly 400 to 450 Elo points, a large margin on a head-to-head scale.
Independent testing shows an 18-point jump on the Intelligence Index
Artificial Analysis confirms the gain on its own standardized tests. At Max effort, Sonnet 5.5 scores 56 on its Intelligence Index, against 38 for Sonnet 5 at Max.
That moved Sonnet 5.5 to second place on the index at launch, behind only Claude Opus 5.5. On the firm's own Terminal-Bench 4.0 run, Sonnet 5.5 scored 64%, about 50 points above Sonnet 5.
These results come with a caveat. The firm tested a pre-release Sonnet 5.5 deployment that had a structured-output bug, since fixed, and plans to rerun the affected evaluations.
Same token price, lower benchmark cost per task at every effort level except Max
On Artificial Analysis's weighted benchmark workload, Sonnet 5.5 is cheaper per task than Sonnet 5 at Low, Medium, High, and Xhigh effort. It also scores 12 to 18 points higher at each of those levels. The table pairs both models' results from the Artificial Analysis release pages.
Table 3 - Sonnet 5.5 vs Sonnet 5 Intelligence Index score and estimated cost per weighted index task by effort level, independently measured by Artificial Analysis, as of September 2026. Figures reflect a benchmark workload, not a production bill.
1. Sonnet 5.5 at Low nearly matches Sonnet 5 at its best for a twelfth of the cost
The most striking row pair is across effort levels. Sonnet 5.5 at Low scores 36 for $0.41 per task, two points short of Sonnet 5's best score of 38 at Max, which costs $5.09.
At Medium, Sonnet 5.5 passes Sonnet 5's best outright, scoring 41 for $0.59. Anthropic's own testing points the same way. It reports that at Medium effort, Sonnet 5.5 beats Sonnet 5's best Terminal-Bench score for less than a tenth of the cost per task.
2. Max effort costs 49% more than Sonnet 5 at Max
Max is the one setting where Sonnet 5.5 costs more. It spends $7.60 per index task against $5.09 for Sonnet 5, because it writes far more output at that level.
At Max, Sonnet 5.5 used about 193k output tokens per index task, the most Artificial Analysis had recorded for any model. That is roughly 60% more than Sonnet 5 used at Max. The extra spend buys 18 more index points, but few workloads need it.
3. Testers report far fewer tokens per job at everyday settings
Customer-reported results in Anthropic's launch post show the other side of the token story. These are private evaluations, not public benchmarks. Balyasny Asset Management saw about 121k tokens per answer across 2,441 finance tasks, against 497k for Sonnet 5.
Slack reported about 14% fewer output tokens on its Slackbot evaluations, and Box reported 2.4 times faster runs with 12% fewer total tokens. These teams reported results from their own workloads rather than Max-effort benchmark runs. Anthropic sums it up as up to 30% lower cost per task than Sonnet 5 for most work.
Suggested read: Claude Sonnet 5.5 Pricing: API Rates, Plans, and What a Task Really Costs
Sonnet 5.5 is faster at every effort level
Anthropic says Sonnet 5.5 generates output more than 30% faster than Sonnet 5. The independent speed data shows a larger gap.
Across effort levels, Artificial Analysis measured Sonnet 5.5 at 85 to 139 output tokens per second, against 55 to 75 for Sonnet 5. The gap runs from 43% faster at High to 86% faster at Medium.
Speed per token still isn't time per task. At Max, Sonnet 5.5's heavier output can offset its faster stream. At Low through High, it is both faster per token and lighter on tokens, so jobs should finish sooner.
Four changes to plan for before you switch
Sonnet 5.5 behaves differently from Sonnet 5 in ways that show up without any code change. Anthropic's docs list these four as the ones most likely to affect an existing app.
1. Effort levels are recalibrated
An effort setting on Sonnet 5.5 doesn't produce the same amount of thinking it did on Sonnet 5. Anthropic advises re-testing rather than carrying a setting over. Its starting points are High for most work and Medium for well-specified agent tasks or latency-sensitive chat.
2. Progress notes between tool calls come back differently
When Sonnet 5.5 pauses between tool calls, longer progress notes come back inside its thinking blocks instead of as plain text. An app that shows those notes to users can go quiet between steps until it is configured to display them. Nothing errors, so this is easy to miss in testing.
3. Some requests fall back to Sonnet 5
Sonnet 5.5 is the first Sonnet model to launch with cyber safeguards. When the API's fallback option is turned on, requests that the safeguards decline as higher-risk cybersecurity work can be retried on Sonnet 5. Routine bug fixing is unaffected. In independent testing, the fallback triggered in about 0.1% of tasks.
4. Shorter prompts can now be cached
The minimum prompt length for caching drops from 1,024 tokens on Sonnet 5 to 512 on Sonnet 5.5. Apps with short, repeated instructions can now cache them. Cached input tokens cost 90% less when reused, while output tokens and cache writes are billed as usual.
Conversation history also carries forward. Sonnet 5.5 can read Sonnet 5's thinking blocks, so a conversation that moves up from Sonnet 5 keeps its earlier reasoning. For teams with custom code, Anthropic also lists five breaking API changes in its migration guide. They cover how up-front thinking is turned off, the removal of forced tool use, rules against editing earlier conversation history, an updated computer-use tool, and which models can act as an advisor.
Sonnet 5 stays available until at least June 2027
There is no deadline forcing the switch. Anthropic's model deprecations page lists Sonnet 5 as active, with retirement not sooner than June 30, 2027.
That leaves time to test before moving production work. A few situations justify staying on Sonnet 5 a little longer:
- Your integration depends on a feature Sonnet 5.5 dropped, such as forced tool use, and the code hasn't been updated yet.
- Your prompts and effort settings are heavily tuned, and you need time to re-test them on the recalibrated effort levels.
- Your workload runs at Max effort under a fixed per-task budget. In that case, test Sonnet 5.5 at Xhigh first, where it scores 52 for $2.74 against Sonnet 5's 38 for $5.09.
Upgrade to Sonnet 5.5 for almost every workload
Sonnet 5.5 is the better model at the same price, and it costs less per task at the settings most teams use. The decision table covers the common cases.
Table 4 - When to move from Sonnet 5 to Sonnet 5.5, based on Anthropic's published data and Artificial Analysis results, as of September 2026.
If Sonnet 5.5 still falls short at practical settings, compare Opus 5.5 with Sonnet 5.5 at Max on your own tasks before committing to either. The Sonnet 5.5 vs Opus 5.5 comparison covers where that line falls.
Suggested read: Claude Opus 5.5 vs Claude Sonnet 5: The Ultimate Comparison
Get Sonnet 5.5's gains without touching an API key
Sonnet 5.5 is a strong default upgrade from Sonnet 5, though not a blind drop-in replacement. It keeps the same price, with much higher scores, faster output, and lower benchmark cost per task at every effort level below Max. Switch on your own schedule before Sonnet 5 retires, re-test your effort settings, and skip Max unless the work truly needs it.
Sonnet 5.5 is now live on Emergent. To add it to the apps you build, you can use your own Anthropic API key or the Universal LLM Key, which skips the separate account and key setup.
Usage is billed through Emergent Credits. When you create a custom agent, you choose the language model it reasons with at setup. That makes Sonnet 5.5 a practical fit for a client portal that drafts support replies or an internal tool that turns spreadsheets into reports.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







