Claude Opus 5 Reviews: What Real Users Actually Think

Claude Opus 5 reviews from real users and testers: the love-hate split, benchmarks, $5/$25 pricing, and whether it is worth switching in 2026.

Bhavyadeep Sinh Rathod
Written by
Bhavyadeep
Priyanka Singh
Reviewed by
Priyanka Singh
Published: 
Aug 12, 2026
0
 min read
Table of Contents

TL;DR

  • Reviewers consistently rank Opus 5 at or near the top, with several calling it the best model available.
  • The loudest complaint is personality: verbose, neurotic, and hard to work with, even from people who rate it highly.
  • It shines as an agent and coder, and struggles as a conversational partner or an orchestrator.
  • Price is unchanged from Opus 4.8 at $5/$25 per million tokens, half of Fable 5.
  • Two measured catches back the complaints: a higher hallucination rate and a very slow first response at max effort.


Claude Opus 5 has earned a rare verdict from the people using it: the best model many have tried, and the one they least enjoy working with. That split runs through nearly every review since the July 24, 2026 launch. Testers rank it at or near the top on capability, then complain in the same breath that it is verbose, over-cautious, and tiring to talk to. Anthropic priced it at $5 per million input tokens and $25 per million output, half of Fable 5's rate, and it took the top spot on the independent Artificial Analysis Intelligence Index.

This roundup pulls together what real reviewers and users are saying about Claude Opus 5, from named engineers and podcasters to launch-week forum threads, and sets each reaction against the benchmark data. The goal is a review grounded in lived experience, not a spec sheet.

Reviewers rate it top-tier on output but dislike using it

The dominant review of Claude Opus 5 is a contradiction: people rate it first and enjoy it least. This is the single most repeated pattern across the launch coverage, and it is worth leading with because it shapes everything else.

The pattern shows up first in the people who tested Opus 5 hands-on. Dan Shipper and Katie Parrott at Every, after a week of testing, called it a hard model to love that argued with instructions and stopped before the work was done. Zvi Mowshowitz, who published one of the most detailed independent write-ups, summed up the community mood plainly: an unusually large number of people strongly dislike talking to it, even while acknowledging it is a strong model.

Not every reaction is mixed. Some testers switched everything over. One engineer quoted in Zvi's roundup, Matt Wigdahl, said he swapped to Opus 5 for everything and now brings in Fable only when he needs the extra size. Another, Will, called it most people's new daily driver and said he does not miss Fable despite having spent real money on it. The praise and the frustration are both genuine, which is exactly why the model is worth a careful look.

Users say it is a great worker and a tiring colleague

The clearest theme in the reviews is that Opus 5 does excellent work but is unpleasant to supervise. Reviewers keep drawing the same line between the artifact it produces and the conversation it puts you through to get there.

The behavior reviewers describe is a model that hesitates to act on its own. Testers report it reaching a reasonable conclusion, then asking whether it should proceed or handing the decision back, even when it has the tools to resolve the question itself. Zvi Mowshowitz frames it as a model that is strong locally but reluctant to run the show, one that often needs another model or a human to keep it moving. The recurring description is neurotic and reliant on reassurance: a tedious colleague and an excellent worker.

The complaint that comes up most often has its own community name. Reviewers and Reddit users call it "Claudeslop," the hedging, apologies, and over-long sentences that make replies tiring to read. One widely shared write-up found "too talkative" was the top-voted reaction to the launch. Zvi describes the same fatigue, noting the model leans on the same verbal tics over and over even though it is clearly smart enough not to.

The frustration is real, but so is the workaround most reviewers land on. The fix people report is to stop fighting the model every message: set a lower effort level, trim the scaffolding, and use an output style that reins in the verbosity. Get out of the conversational loop, hand it a clear brief, and judge the result.

Where reviewers agree it excels: agentic work and front-end builds

The strongest praise for Opus 5 is consistent and specific: it is an outstanding agent and coder, especially on well-defined, self-contained tasks. Reviewers who dislike chatting with it still reach for it here.

Reviewers single out front-end and prototyping work as a standout. The pattern they describe is a model that produces detailed, complete builds rather than a plausible shell, best handed a full brief and run asynchronously rather than steered turn by turn. Ethan Mollick reported the same shape: strong on shorter, well-defined tasks, sometimes less ambitious on very long ones.

The benchmark data backs the agent praise squarely. On Artificial Analysis's AA-Briefcase test for long-horizon knowledge work, Opus 5 scores 1720 Elo at max effort, 146 points ahead of Fable 5 while costing about 20% less per task, and takes the top three positions across its highest effort settings. Coding reactions match too. Anthropic's Boris Cherny highlighted its resistance to prompt injection, and engineer Adam Wolff said Opus 5 at medium effort is his go-to for cranking out a pull request.

Benchmark Opus 5 Opus 4.8 Fable 5 Source
Intelligence Index 61 56 60 Artificial Analysis
SWE-bench Pro 79.2% 69.2% 80% Anthropic
AA-Briefcase (Elo, max) 1720 lower 1574 Artificial Analysis
ARC-AGI-3 30.2% 1.5% not published Anthropic

The number reviewers singled out most is ARC-AGI-3, a test of novel problem-solving, where Opus 5 scores roughly three times the next-best model. Zvi and others flag it as the result to watch, while cautioning that some testers found the gain did not transfer cleanly to fully novel held-out puzzles.

The complaint that matters most: users say it is confidently wrong

Several reviewers report a trust problem, and it is the criticism most likely to affect real work. The praise for Opus 5's intelligence sits alongside a recurring warning that it states wrong things with confidence, then folds when challenged.

Engineer Max Weinbach, quoted in Zvi's roundup, described switching back to GPT-5.6 Sol after Opus 5 confidently claimed it had not done something it had in fact done correctly. His framing was precise: the issue was not that Opus is bad, it was that he could not trust it. Others echoed a tendency to confuse what the model said with what the user said.

The measured data supports the complaint. Artificial Analysis found Opus 5's hallucination rate rose on its closed-book factual benchmark even as raw accuracy improved, reaching roughly 50%, and Anthropic's own system card records a similar trade. The mechanism is simple: the model answers more often when it is uncertain. On a coding task with tests, a wrong guess is cheap because the test catches it. On a factual claim in a report, a confident wrong guess is the failure itself.

Reviewers split on whether it beats Opus 4.8

Ask whether Opus 5 is actually better than the model it replaces, and the reviews divide cleanly. On the benchmarks the answer is not close: Opus 5 beats Opus 4.8 across the board at the same price. On lived experience, users split. That gap between the numbers and the feel is the whole story of this section.

Start with what the benchmarks say. Every published head-to-head favors Opus 5, often by a wide margin, and the two models cost the same $5/$25 per million tokens.

Benchmark Opus 5 Opus 4.8 What it measures
Intelligence Index 61 56 Aggregate capability
SWE-bench Pro 79.2% 69.2% Contamination-resistant coding
SWE-bench Multimodal 59.4% 38.4% Coding from visual input
AutomationBench 26.0% 17.0% Business workflow automation
ARC-AGI-3 30.2% 1.5% Novel problem-solving

Now set that against what users report, because this is where the split appears. On Hacker News, one user reported it looks worse than Opus 4.8 on some tasks, saying it needed more hand-holding on an existing codebase before it stopped doing redundant work. Another, after six hours of use, called it absurd to say it is worse and described it as wonderful. Both are describing the same model, which points at the real cause rather than a regression.

Two changes explain the split. The first is output drift: Opus 5 answers the same prompt differently from Opus 4.8, in tone, formatting, and edge-case handling, so prompts tuned for the older model can read as worse. The second is that extended thinking is on by default and scales with task difficulty, which makes answer depth and length vary more from request to request. Reviewers who tuned their setup for Opus 4.8 felt the friction most; those who started fresh felt it least.

Also read our Claude Opus 5 vs Opus 4.8 breakdown for a full side-by-side before you decide which one fits your workflow.

How to set the effort dial, according to the people using it

Reviewers converge on one practical setting: do not default to max effort. This is the most actionable advice in the entire body of Opus 5 coverage, and it emerged from users noticing the model behaving oddly at high effort.

The pattern reviewers spotted is counterintuitive. Zvi and several Hacker News users found Opus 5's coding score can fall at effort levels above medium. Anthropic's system card explains why: at higher effort the model makes more changes than the task requires, and a grader penalizes the extra work as out of scope. A short "stay in scope" instruction recovers most of the lost performance. As one tester put it, the model goes down too many rabbit holes and over-engineers without being asked.

The effort setting runs from low through medium, high, xhigh, and max. Use this as a rough guide:

Low or medium: routine, high-volume, or latency-sensitive calls, and the level several reviewers prefer for everyday coding.

High: the best all-round value point for hard reasoning and agentic work, and Anthropic's recommended starting level.

Xhigh or max: reserve for genuinely difficult single steps, and expect the highest token spend and the slowest responses.

That slowness is not just felt. Artificial Analysis measured time to first token at about 68 seconds at max effort, against a class median under three seconds, which is why reviewers who run interactive sessions keep the effort low.

Reviewers treat it as a subagent, not the boss

A specific and repeated verdict is that Opus 5 works best taking orders, not giving them. Reviewers describe it as an excellent executor that can lose the thread when asked to run a whole project.

Zvi's framing captures the consensus: Opus 5 is very good at thinking locally and less good at thinking globally, an excellent subagent that sometimes needs another model or a human to keep it on track. Several engineers in his roundup reported using Fable 5 as an orchestrator directing Opus 5 subagents, and one noted that Fable consistently reports Opus 5 produces good code, including catching spec errors in badly written requests. The recurring picture is a strong pair of hands that benefits from an adult in the room on long, open-ended runs.

Opus 5 vs Fable 5, GPT-5.6 Sol, and Kimi K3: what reviewers recommend

The consensus is that Opus 5 is the value pick and Fable 5 is the ceiling, with the capability gap smaller than the price gap. On the Artificial Analysis Intelligence Index, Opus 5 scores 61 to Fable 5's 60, so the cheaper model narrowly leads on measured intelligence, while Fable 5 still wins the hardest reasoning and remains the frontier flagship.

Reviewers describe the difference in feel more than in scores. Theo of t3.gg called Opus 5 a useful in-between of GPT-5.6 Sol and Fable 5, with Fable's taste and Sol's thoroughness, writing code slightly less pretty than Fable's but more likely to be correct. Others called it "Fable Lite," strong but a touch less creative or big-picture. Against the wider field, the top is close: on the same index, GPT-5.6 Sol scores 59 and Kimi K3 scores 57, so three leading models sit within a few points, and Opus 5's clearest separation is on agentic work rather than raw intelligence.

Model Intelligence Index Price (input / output) Reviewers reach for it when
Claude Opus 5 61 $5 / $25 Agentic tasks, coding, value
Claude Fable 5 60 $10 / $50 The hardest reasoning, chat
GPT-5.6 Sol 59 lower per token Direct answers, workhorse tasks
Kimi K3 57 lower per token Budget-sensitive work

The practical routing most reviewers land on: Opus 5 for well-defined coding and agent tasks, Fable 5 when you need maximum intelligence or a nicer conversation, and a cheaper model for short or latency-sensitive work.

Want the full breakdown of what Opus 5 costs across plans and use cases? Read our Claude Opus 5 pricing guide before you commit.

Who should switch, based on the reviews

Switch to Opus 5 if you run agents or write code, and stay put if your work is short, latency-sensitive chat. The reviews divide cleanly by workload rather than by taste.

The people happiest with Opus 5 run long multi-step tasks, coding loops, and prototyping, where its self-verification and low price make it a straightforward upgrade from Opus 4.8. Current Fable 5 users report switching much of their workload over and keeping Fable for the hardest jobs. The people most frustrated tried to use it as a conversational partner or dropped it into a workflow tuned for an older model, and hit the verbosity and the eager-to-act behavior head on. If that is your use, a cheaper, faster model like Sonnet 5 is often the better call.

Beyond the model comparison

Picking the right model is one decision, but shipping the product that runs on it is the harder one. The reviews above agree on a deeper point: a model is only as useful as the workflow around it, and even the best model still needs a backend, integrations, and real code before it becomes something people can use.

That is the gap Emergent is built to close. Emergent lets you describe the app you want and builds a full-stack, production-ready product with a real backend, real integrations, and code you own. It runs on leading models including Claude, OpenAI GPT, and Google Gemini through the Universal LLM Key, a single credential that reaches every supported model with unified billing through Emergent Credits and no separate API setup, so you can switch between models without reconfiguring anything.  

Start Building to turn your idea into a live app.

Was this article helpful?
About the writer
Bhavyadeep
Bhavyadeep Sinh Rathod
Content Manager

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free

Frequently Asked Questions

Your Questions, Answered

Is Claude Opus 5 worth it over Opus 4.8?
Most reviewers say yes for agent and coding work. Opus 5 costs the same per token as Opus 4.8 and adds stronger reasoning and better long-agent reliability. The main caveat from users is output drift, so test your own prompts against both models before switching a tuned workflow, since answer style can shift.
Which is the best Opus model?
Opus 5 is the best Opus model currently available and the default on Claude Max. It replaces Opus 4.8, which becomes legacy. Reviewers rank it at the top of the Opus line and near the top of all models, and it keeps the same price as its predecessor, so for most Opus workloads it is the clear choice.
Can I use Claude Opus 5 for free?
No. Opus 5 is not on the Claude Free plan, which gives you Sonnet 5 instead. You can access Opus 5 on Claude Pro, Max, Team, and Enterprise plans, and through the API. It is the default model on Claude Max and the strongest model available to Pro subscribers.
Why do reviewers say Opus 5 is annoying?
The most common complaints are verbosity, described by users as "Claudeslop," and a cautious, over-eager personality that asks for confirmation or acts before you are ready. Even reviewers who rank it first mention this. The fix people recommend is a lower effort setting, trimmed instructions, and an output style that constrains the tone.
Is Claude Opus 5 better than Fable 5?
On the Artificial Analysis Intelligence Index, Opus 5 edges Fable 5, 61 to 60, at half the price, and reviewers agree it matches Fable on most everyday work. Fable 5 still leads the hardest reasoning and is the more pleasant conversational model, so many reviewers keep it for chat and use Opus 5 for agentic tasks.
Is Claude Opus 5 good for coding?
Yes, and coding is where reviewer praise is strongest. Engineers report using it as their default coder, often at medium effort, and Fable 5 has been observed to rate Opus 5's code well. It more than doubles Opus 4.8 on hard coding benchmarks. Watch the safety-classifier fallback on its headline coding score and verify on your own repository before migrating.
Start Building
on Emergent today
Try Emergent
This is some text inside of a div block.
This is some text inside of a div block.
Note

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

https://api.linear.app/graphql