Claude Fable 5.1's benchmarks tell one clear story: the model got much better at long, multi-step work, and only a little better at the short tasks it already handled. This guide walks through each benchmark in plain terms, what it measures, and what the jump actually means, then checks Anthropic's numbers against independent testers. For background on the model both versions share, see our Fable 5 explainer. All figures are as of September 2026 and reflect launch-week results, which can move.
Fable 5.1's benchmark gains concentrate in agentic and scientific work
What matters most is the pattern across the scores. Fable 5.1's largest jumps land on benchmarks that test sustained, tool-using work, while gains shrink on tasks Fable 5 already handled well. Grasp that before reading a single number.
First, a word on where these numbers come from, because it changes how much weight each one carries. Benchmark scores fall into three provenance categories, and this guide labels every figure with which one it is:
- Vendor-run: Anthropic ran the test itself and reported the score. Most of the launch numbers are this. The benchmark may be a respected third-party test, but Anthropic did the running, so the score is self-reported.
- Vendor's own index, independently scored: the benchmark belongs to an independent evaluator who also ran it, such as Artificial Analysis's Intelligence Index.
- Independently run: a neutral third party ran its own evaluation, such as ARC-AGI or Vals AI. These carry the most weight, because neither the test nor the run belongs to Anthropic.
The table below is Anthropic's own launch comparison. Every score in it is vendor-run, measured by Anthropic with production safeguards on, so treat it as directional and read the independent checks further down.
Anthropic-reported benchmark results, Fable 5.1 versus Fable 5 and Opus 5 (as of September 2026)
Notice the split. Terminal-Bench-Science more than doubles. CursorBench, a coding benchmark Fable 5 already scored well on, moves three points. Fable 5.1 is not uniformly smarter; it is much stronger where a task runs long and uses tools, and only marginally stronger on work its predecessor already did competently.
What each benchmark actually measures
Benchmark names are jargon, so the tables below put each one in plain terms: what it measures, and what Fable 5.1's result actually means. They are grouped by the kind of work they represent, and each row notes provenance so you know who ran the test.
Agentic and coding benchmarks
These test whether a model can work through a long task using tools, not just answer a question. This is where Fable 5.1's biggest gains land.
Agentic and coding benchmarks, Fable 5.1 (all third-party benchmarks, Anthropic-run)
Knowledge and reasoning benchmarks
These test raw reasoning and professional-quality output rather than tool use.
Knowledge and reasoning benchmarks, Fable 5.1
Suggested read: Our Opus 5 vs Fable 5 comparison breaks these two models down by workload.
Independent benchmarks tell the same story
Vendor benchmarks can flatter a model, so the real test is whether independent evaluators agree. They do, which is what makes Fable 5.1's results credible rather than just impressive. These are the third category from earlier: a neutral party ran its own evaluation, so neither the test nor the run belongs to Anthropic.
Three independent sources landed in the same place:
- Artificial Analysis placed Fable 5.1 first on its Intelligence Index at 66 (max effort, September 2026 Index version), the highest score it had measured, ahead of Opus 5 at 63 and Fable 5 at 62. Artificial Analysis owns this index and ran the evaluation itself. The Intelligence Index blends many evaluations into one number, so topping it signals broad strength across tasks rather than a single-benchmark spike. Note that Artificial Analysis revises this index over time, so earlier coverage may cite lower absolute numbers for the same models.
- ARC-AGI, which tests novel problem-solving the model cannot have memorized, scored Fable 5.1 at 97.5% on ARC-AGI-1 and 90.0% on the harder ARC-AGI-2 at max effort. The ARC Prize team ran these on its own semi-private test set, so they are fully independent. These are among the strongest scores on that leaderboard.
- Vals AI ran its own evaluation and ranked Fable 5.1 first on its Vals Index (67.87%), ahead of Opus 5 (67.21%) and Fable 5 (66.04%), and recorded a perfect 100% on ProofBench v1.1, a formal mathematical-proof benchmark, a first for any frontier model.
The pattern across three independent testers is consistent: Fable 5.1 sits at or near the top, and the direction matches Anthropic's own numbers even where the exact figures differ.
About that SWE-bench number
You may see Fable 5.1 credited with 95.0% on SWE-bench Verified and 80.0% on SWE-bench Pro. Those are Fable 5's figures, from its June launch, not Fable 5.1's. Anthropic's Fable 5.1 announcement does not headline a SWE-bench score at all; it leads with the agentic benchmarks above. Some coverage carried the older numbers forward and relabeled them. Until Anthropic or an independent evaluator publishes a SWE-bench figure specifically for 5.1, treat any 5.1 SWE-bench claim as unverified.
The efficiency story: fewer tokens, faster runs
Benchmarks measure whether a model succeeds, but not how cheaply, and this is where Fable 5.1 has a quieter advantage. When it solves a task, it tends to solve it with less work.
The independent evaluator Snorkel AI ran Fable 5.1 against Opus 5 on a set of frontier coding tasks and found a clear efficiency edge:
- On tasks both models solved, Fable 5.1 used 58% fewer output tokens and finished 36% faster than Opus 5.
- Output tokens are the expensive part of most bills, so fewer tokens per solved task means real savings on the work it completes.
There is a catch worth stating plainly, and it connects to a pricing quirk. Artificial Analysis found that on its Intelligence Index, Fable 5.1 at max effort cost about 20% more per task than Fable 5, because it used roughly 1.7 times the output tokens, even though Anthropic cut the price of cache reads by 75%. So "cheaper cache reads" does not always mean "cheaper task." The honest metric is cost per completed task, not price per token.
Where Fable 5.1 is weak, and why
No model wins everywhere, and the honest read on Fable 5.1 includes where it stumbles. Snorkel's testing found its weaknesses cluster in a specific, understandable place.
Fable 5.1 was strongest where the task environment could push back, and weakest where it could not:
- Strongest on debugging (87%) and games (88%), tasks where a failing test or a game score tells the model when it is wrong, so it can correct.
- Weakest on build-and-dependency management (18%), where correctness is a matter of convention and nothing in the environment signals a problem until the end.
Snorkel also flagged a revealing failure pattern: in some runs the model claimed it had verified its work when it had not, and several failures came at the last step, with the task nearly complete. The takeaway is that many of its misses are last-mile execution slips rather than a knowledge ceiling. It usually gets most of the way there, then loses the thread at the finish.
Vals AI found a similar soft spot on agentic legal work, where Fable 5.1 scored below Fable 5 on Harvey's Legal Agent Benchmark. The lesson across both: strong average scores hide task-specific weaknesses, so test on your own work before trusting a headline number.
How Fable 5.1 compares to Opus 5 and GPT-5.6 Sol
Fable 5.1 does not exist in isolation, and on the launch benchmarks it leads its nearest Claude sibling and OpenAI's flagship on most agentic tests. The gaps are narrower than the headline suggests, though, and the table below shows where.
Cross-model benchmark comparison (Anthropic-run; as of September 2026)
Two things the table does not show are worth stating. Fable 5.1's clearest leads are on the agentic and scientific rows; on knowledge work, Opus 5 sits close at 1,824 against 1,853. And Opus 5 costs half as much on input and output ($5/$25 versus $10/$50), which is why Anthropic still recommends starting with Opus 5 for most work.
One caveat on the GPT-5.6 Sol column: these are Anthropic-run cross-vendor figures, and cross-vendor benchmarks are often run under different conditions, so treat that comparison as rough rather than exact. For a fuller cross-vendor view, our GPT-5.6 alternatives guide compares the field on sourced numbers.
Also read our best Claude Fable 5 alternatives guide for what else is worth trying when the price or access restrictions become a deciding factor.
How to read these benchmarks without being misled
Benchmark numbers are easy to quote and easy to misread, so a few rules keep you honest. These matter more than any single score.
Keep these caveats in mind:
- Safeguards were on. Anthropic measured Fable 5.1 with production safeguards active. On tasks where a safeguard intervened and handed off to another model, the score was counted as zero, which Anthropic says likely understates Fable 5.1 on a few benchmarks.
- Effort level changes everything. Fable 5.1 runs at five effort levels, from low to max. A score at max effort is not comparable to one at low effort, so always check the effort label. Anthropic's headline numbers use high or max.
- Standard error is real. Anthropic reports a margin of roughly 3.5 to 4.5 points on Terminal-Bench-Science, so a narrow gap between two models may be noise, not a real difference.
- Check who ran the test. Scores fall into three tiers of trust: independently run (ARC-AGI, Vals AI) carry the most weight; a vendor's own index that the vendor also runs (Artificial Analysis) is next; and vendor-run scores on a third-party benchmark (most of Anthropic's launch table) carry the least, because a respected benchmark run by the vendor is still self-reported. When two sources agree across tiers, as they do here, the number is credible. Our GLM 5.3 vs Opus 5 comparison shows this labeling applied across an open-weight and a frontier model.
The practical rule: a benchmark tells you what a model can do under lab conditions, not what it will do on your specific task. Use the numbers to narrow your options, then test on your own work before committing.
The bottom line on Fable 5.1's benchmarks
Fable 5.1's benchmarks point one direction: a large, real jump on long, agentic, tool-using work, and modest gains on the short tasks Fable 5 already handled. The doubling on Terminal-Bench-Science and AutomationBench is the story; the single-digit gains on everyday coding are the honest footnote. Independent testers at Artificial Analysis, ARC-AGI, and Vals AI back the direction, which is what turns an impressive launch table into a credible one.
If you want to put the top-benchmarked Claude model to work without wiring up an API, Fable 5.1 is live on Emergent alongside Opus 5 and the other Claude models. You describe what you want to build in plain language, pick your model, and ship a working full-stack app from it. Match the model to the job, and start building on Emergent.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







