Models · Anthropic · Out sinceReleased 1 Sep 2026
Claude Fable 5.1#4 for writing.#4 for writing, best at max effort.
Claude Fable 5.1 is made by Anthropic. Among the models we track it ranks #4 for writing, #5 for research and analysis, #5 for coding. It's expensive to use. When an AI has to work on its own for a long time — using a computer, browsing, or handling a whole task end-to-end — this one is the most reliable.
API id claude-fable-5-1. Same model as Claude Mythos 5.1 with stricter safeguards (fallback to Opus models on cyber/bio). Default effort high (high in Claude Code, medium in Cowork/claude.ai). Cache reads $0.25/M (75% cheaper than Fable 5).
- 74.9 / 100 · #4
- 80.1 / 100 · #5
- 81.8 / 100 · #5
- Expensive$10 / $50
- $20
- 1M
- 128K
- 205 (135 independent135 indep.)
- 1 Sep 2026
- Anthropic
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (Maximum thinking).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-fable-5.1.
Route it as anthropic/claude-fable-5.1 at $10 in / $50 out per 1M tokens, 1M context. Listed since 1 Sep 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstool_choicetoolsverbosity
Thinking level
Reasoning effort
How long should Claude Fable 5.1Where Claude Fable 5.1think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Artificial Analysis Coding Agent Index
Best at MaxBest at Max
Terminal-Bench 4.0
Best at Extra highBest at Extra high
CursorBench
Best at MaxBest at Max
DeepSWE
Best at MaxBest at Max
FrontierCode
Best at Medium, worse abovePeaks at Medium
CursorBench 3.2.0
Best at MaxBest at Max
SWE Atlas - Codebase QnA
Best at Extra high, worse abovePeaks at Extra high
Terminal-Bench 2.1
Best at MaxBest at Max
SciCode
Best at MaxBest at Max
SWE-bench Pro (Anthropic internal subset)
Best at HighBest at High
Artificial Analysis Intelligence Index
Best at MaxBest at Max
Terminal-Bench-Science 0.1
Best at MaxBest at Max
AA-Briefcase v1.1
Best at MaxBest at Max
AA-LCR
Best at MaxBest at Max
AA-Omniscience Accuracy
Best at MaxBest at Max
AA-Omniscience Hallucination Rate
Best at Low, worse abovePeaks at Low
Lower is better on this test
AA-Omniscience Index
Best at MaxBest at Max
ARC-AGI-2
Best at Extra highBest at Extra high
GDPval-AA v2.1
Best at MaxBest at Max
GPQA Diamond
Best at MaxBest at Max
Harvey LAB-AA
Best at Extra high, worse abovePeaks at Extra high
Humanity's Last Exam
Best at MaxBest at Max
Humanity's Last Exam (with tools)
Best at Extra high, worse abovePeaks at Extra high
WANDR
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA Analyst Agent | 57.5% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run; scores move in 1.25-pt steps (small task set) |
| AA-Briefcase | 1694 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | |
| AA-Briefcase v1.1 | 1669 | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Briefcase v1.1 | 1678 | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Briefcase v1.1 | 1678 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| AA-LCR | 82.3% | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 84.7% | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 83.7% | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 83.0% | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 85.3% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 60.2% | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 63.1% | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 64.9% | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 66.2% | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 67.2% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 65.6% | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 69.1% | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 68.8% | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 70.5% | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 72.6% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 34.1 | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 37.6 | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 40.8 | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 42.4 | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 43.5 | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| ARC-AGI-1 | 97.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | |
| ARC-AGI-2 | 86.2% | MediumClaude Fable 5.1 (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $1.22/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 88.8% | HighClaude Fable 5.1 (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $1.67/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 90.0% | Extra highClaude Fable 5.1 (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $3.12/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 90.0% | MaxClaude Fable 5.1 (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $4.49/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 90.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | |
| Artificial Analysis Coding Agent Index | 61.7 | Extra highClaude Fable 5.1 XHigh + SWE-2 Medium | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Devin Fusion CLI | agent Devin Fusion CLI; components: DeepSWE v1.1 63.1, SWE-Atlas-QnA 65.9, Terminal-Bench v4 56.1; avg cost $7.90/task; avg wall time 36 min/task; Devin Fusion = Fable/GPT planner + Cognition SWE-2 (medium) sidekick |
| Artificial Analysis Coding Agent Index | 62.2 | MaxFable 5.1 (max) (with fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | agent Claude Code; components: DeepSWE v1.1 64.3, SWE-Atlas-QnA 64.8, Terminal-Bench v4 57.6; avg cost $12.39/task; avg wall time 35 min/task; with fallback model |
| Artificial Analysis Intelligence Index | 46.8 | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-fable-5-1-low; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $2.37/task |
| Artificial Analysis Intelligence Index | 48.9 | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-fable-5-1-medium; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $2.98/task |
| Artificial Analysis Intelligence Index | 51.2 | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-fable-5-1-high; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $3.91/task |
| Artificial Analysis Intelligence Index | 53.2 | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-fable-5-1-xhigh; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $5.98/task |
| Artificial Analysis Intelligence Index | 53.4 | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-fable-5-1; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $7.63/task |
| Artificial Analysis output speed | 47 tok/s | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 4.5s; list price $10/50 per 1M in/out |
| Artificial Analysis output speed | 49 tok/s | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 8.1s; list price $10/50 per 1M in/out |
| Artificial Analysis output speed | 51 tok/s | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 29.7s; list price $10/50 per 1M in/out |
| Artificial Analysis output speed | 59 tok/s | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 102.1s; list price $10/50 per 1M in/out |
| Artificial Analysis output speed | 68 tok/s | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 262.3s; list price $10/50 per 1M in/out |
| ArXivMath (no tools) | 82.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | Aug 2026 |
| ArXivMath (with tools) | 92.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | Aug 2026 |
| AutomationBench | 31.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps) |
| Chartography (with tools) | 88.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| CursorBench | 45.1% | LowFable 5.1 Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 15; $5.44/task; 34,795 tokens/task; 51 steps/task |
| CursorBench | 45.1% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $5.44 |
| CursorBench | 46.8% | MediumFable 5.1 Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 11; $7.05/task; 45,411 tokens/task; 63 steps/task |
| CursorBench | 46.8% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $7.05 |
| CursorBench | 49.2% | HighFable 5.1 High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 9; $9.08/task; 58,438 tokens/task; 77 steps/task |
| CursorBench | 49.2% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $9.08 |
| CursorBench | 51.6% | Extra highFable 5.1 Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 8; $13.01/task; 87,294 tokens/task; 101 steps/task |
| CursorBench | 51.6% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $13.01 |
| CursorBench | 51.8% | MaxFable 5.1 Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 7; $17.28/task; 117,236 tokens/task; 128 steps/task |
| CursorBench | 51.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $17.28 |
| CursorBench 3.2.0 | 66.2% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Cursor agent | CursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $2.9 |
| CursorBench 3.2.0 | 68.0% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Cursor agent | CursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $3.53 |
| CursorBench 3.2.0 | 69.4% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Cursor agent | CursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $4.8 |
| CursorBench 3.2.0 | 72.8% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Cursor agent | CursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $6.96 |
| CursorBench 3.2.0 | 73.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Cursor agent | CursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $9.64 |
| DeepSWE | 63.1% | Extra highClaude Fable 5.1 XHigh + SWE-2 Medium | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Devin Fusion CLI | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 64.3% | MaxFable 5.1 (max) (with fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 67.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | |
| Design Arena (all categories) | 1342 | Defaultclaude-fable-5-1 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±3.8 SE; 9335 battles; win rate 59.3% |
| Design Arena (fullstack) | 1324 | Defaultclaude-fable-5-1 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±15.9 SE; 509 battles; win rate 57.2% |
| Epoch Capabilities Index | 164.8 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 161.8-169.1; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 2162 | Default*claude-fable-5-1 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 2; rubric score 16.95/20; slop 8.16; avg length 5841 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*) |
| FrontierCode | 49.8% | Loweffort=low | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $2.38; competitor number as plotted by OpenAI |
| FrontierCode | 52.8% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Fable 5.1 adds unrequested out-of-scope edits at higher effort; vendor-reported cost/task $2.47 |
| FrontierCode | 50.9% | Mediumclaude-fable-5-1_medium | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | Claude Code | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5 |
| FrontierCode | 50.9% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Fable 5.1 adds unrequested out-of-scope edits at higher effort; vendor-reported cost/task $3.28 |
| FrontierCode | 50.3% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Fable 5.1 adds unrequested out-of-scope edits at higher effort; vendor-reported cost/task $5.27 |
| FrontierCode | 48.7% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Fable 5.1 adds unrequested out-of-scope edits at higher effort; vendor-reported cost/task $9.27 |
| FrontierCode | 50.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Fable 5.1 adds unrequested out-of-scope edits at higher effort; vendor-reported cost/task $12.82 |
| FrontierCode v1.1 (Extended) | 63.6% | Mediumadaptive thinking, effort=medium (best) | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Claude Code | Fable 5.1 peaks at medium on FrontierCode; adds unrequested out-of-scope edits at higher effort |
| FrontierMath (Tiers 1-3) | 90.2% | Maxclaude-fable-5-1_max | Independent testIndependentEpoch AI ↗ | 1 Sep 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 4 | 87.8% | Maxclaude-fable-5-1_max | Independent testIndependentEpoch AI ↗ | 1 Sep 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 5.2pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierSWE v2 | 56.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Proximal harness | reported as 0.57 on a 0-1 scale (Opus 5.5 card later lists 56.3%) |
| GDPval-AA v2 | 1853 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | GDPval-AA v2 Elo |
| GDPval-AA v2.1 | 1450 | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $1.41 |
| GDPval-AA v2.1 | 1536 | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $2.17 |
| GDPval-AA v2.1 | 1617 | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $3.43 |
| GDPval-AA v2.1 | 1721 | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1695.77-1746.08 |
| GDPval-AA v2.1 | 1721 | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $7.09 |
| GDPval-AA v2.1 | 1735 | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1710.09-1759.25 |
| GDPval-AA v2.1 | 1735 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $9.59 |
| GPQA Diamond | 88.1% | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-low) |
| GPQA Diamond | 88.6% | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-medium) |
| GPQA Diamond | 90.6% | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-high) |
| GPQA Diamond | 93.4% | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-xhigh) |
| GPQA Diamond | 93.7% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1) |
| GPQA Diamond | 93.4% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.863 stderr; $0.205/test |
| GSO | 88.2% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | 27 Sep 2026 | OpenHands | Opt@1; hack-controlled score 87.25; data https://gso-bench.github.io/assets/leaderboard.json |
| Harvey LAB-AA | 93.0% | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| Harvey LAB-AA | 93.3% | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| Harvey LAB-AA | 93.0% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| HealthBench Professional | 62.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | |
| HLE Diamond | 51.3% | Defaultclaude-fable-5-1 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 3 (Scale rank accounts for CI); ±3.1; entry added 2026-04-10 |
| Humanity's Last Exam | 48.9% | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-low) |
| Humanity's Last Exam | 53.2% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | Humanity's Last Exam without tools; vendor-reported cost/task $0.3 |
| Humanity's Last Exam | 53.8% | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-medium) |
| Humanity's Last Exam | 55.9% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | Humanity's Last Exam without tools; vendor-reported cost/task $0.46 |
| Humanity's Last Exam | 55.9% | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-high) |
| Humanity's Last Exam | 58.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | Humanity's Last Exam without tools; vendor-reported cost/task $0.75 |
| Humanity's Last Exam | 58.7% | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-xhigh) |
| Humanity's Last Exam | 46.5% | Extra highFable 5.1 (xhigh) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±2; entry added 2026-09-03 |
| Humanity's Last Exam | 60.4% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | Humanity's Last Exam without tools; vendor-reported cost/task $1.53 |
| Humanity's Last Exam | 59.1% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1) |
| Humanity's Last Exam | 60.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | Humanity's Last Exam without tools; vendor-reported cost/task $2.23 |
| Humanity's Last Exam (with tools) | 60.0% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | Humanity's Last Exam with tools (web search/fetch, code execution); vendor-reported cost/task $0.52 |
| Humanity's Last Exam (with tools) | 63.0% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | Humanity's Last Exam with tools (web search/fetch, code execution); vendor-reported cost/task $0.67 |
| Humanity's Last Exam (with tools) | 64.8% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | Humanity's Last Exam with tools (web search/fetch, code execution); vendor-reported cost/task $1.05 |
| Humanity's Last Exam (with tools) | 65.1% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | Humanity's Last Exam with tools (web search/fetch, code execution); vendor-reported cost/task $2.28 |
| Humanity's Last Exam (with tools) | 65.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | Humanity's Last Exam with tools (web search/fetch, code execution); vendor-reported cost/task $3.2 |
| IOI (Vals) | 90.8% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.647 stderr; $11.197/test |
| LegalBench (Vals) | 88.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-fable-5-1; rank 2/149; ±0.42 stderr; $0.051606/test |
| LiveCodeBench | 90.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.863 stderr; $0.261/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 3.8 | HighClaude Fable 5.1 (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 1/56; Thurstone comparison score (centered at 0); est. win chance 93%; 95% bootstrap 3.702 to 3.896 |
| LMArena Code Arena (WebDev) | 1751 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 4 (CI rank 3-4); 95% CI 1741-1760; 6137 votes |
| LMArena Text - Coding category | 1533 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 17 (CI rank 2-48); 95% CI 1520-1545; 2254 votes |
| LMArena Text - Creative Writing | 1482 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 10 (rank range 4-26); 95% CI 1468.8-1494.5; 2506 votes; style-controlled |
| LMArena Text - Expert | 1519 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 16 (rank range 3-62); 95% CI 1499.9-1538.6; 965 votes; style-controlled |
| LMArena Text - Hard Prompts | 1522 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 7 (rank range 2-19); 95% CI 1513.6-1529.4; 6630 votes; style-controlled |
| LMArena Text - Instruction Following | 1498 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 7 (rank range 3-20); 95% CI 1486.9-1508.5; 3380 votes; style-controlled |
| LMArena Text - Longer Query | 1511 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 7 (rank range 2-21); 95% CI 1502.0-1520.8; 4705 votes; style-controlled |
| LMArena Text - Multi-Turn | 1487 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 25 (rank range 7-74); 95% CI 1471.1-1502.7; 1474 votes; style-controlled |
| LMArena Text - Non-English | 1495 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 3 (rank range 1-13); 95% CI 1487.6-1503.2; 6689 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1487 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 15 (rank range 3-48); 95% CI 1473.7-1499.8; 2102 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1494 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 19 (rank range 1-74); 95% CI 1474.2-1514.4; 918 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1493 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 7 (rank range 2-19); 95% CI 1481.5-1504.0; 3159 votes; style-controlled |
| LMArena Text (overall) | 1501 | Maxclaude-fable-5.1-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 6 (CI rank 2-14); 95% CI 1495-1508; 11241 votes |
| MCP Atlas | 87.2% | DefaultFable 5.1 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 1 (Scale rank accounts for CI); ±2.05; entry added 2026-09-04 |
| OfficeQA Pro | 69.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| OSWorld 2.0 | 77.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | partial score; August 2026 task release; safeguard interventions scored 0 |
| OSWorld 2.0 (strict pass rate) | 41.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | strict pass rate; August 2026 task release |
| OSWorld 2.1 (partial score) | 80.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | partial score; re-evaluated by Anthropic under Opus 5.5 conditions |
| OSWorld 2.1 (strict pass rate) | 42.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| PRBench Finance (Scale) | 50.8% | DefaultFable 5.1 | Independent testIndependentScale AI (SEAL) ↗ | 3 Sep 2026 | — | Scale rank 7; ±0.22 CI |
| PRBench Legal (Scale) | 51.6% | DefaultFable 5.1 | Independent testIndependentScale AI (SEAL) ↗ | 3 Sep 2026 | — | Scale rank 5; ±0.24 CI |
| ProgramBench (fully resolved) | 7.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.809 stderr; $58.517/test; strict fully-resolved rate |
| ProgramBench (fully resolved) | 87.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | mini-swe-agent | |
| SciCode | 56.7% | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-low) |
| SciCode | 56.4% | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-medium) |
| SciCode | 58.7% | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-high) |
| SciCode | 60.9% | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-xhigh) |
| SciCode | 63.1% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1) |
| SimpleBench | 86.6% | DefaultClaude Fable 5.1 | Independent testIndependentSimpleBench ↗ | 3 Sep 2026 | — | AVG@5, temp 0.7; rank 2nd; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 70.8% | Maxclaude-fable-5-1_max | Independent testIndependentEpoch AI ↗ | 1 Sep 2026 | — | Epoch-run (no tools); ±1.44 stderr |
| SkillsBench | 61.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.76 stderr; $4.517/test |
| SWE Atlas - Codebase QnA | 65.9% | Extra highClaude Fable 5.1 XHigh + SWE-2 Medium | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Devin Fusion CLI | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 60.0% | Extra highFable-5.1 (Claude Code) xHigh* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 1 (Scale rank accounts for CI); ±4.85; entry added 2026-03-02; * = Scale footnote (see page) |
| SWE Atlas - Codebase QnA | 64.8% | MaxFable 5.1 (max) (with fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Refactoring | 56.7% | Extra highFable-5.1 (Claude Code) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 1 (Scale rank accounts for CI); ±6.52; entry added 2026-09-02 |
| SWE Atlas - Test Writing | 67.0% | Extra highFable-5.1 (Claude Code) xHigh* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 1 (Scale rank accounts for CI); ±5.33; entry added 2026-09-02; * = Scale footnote (see page) |
| SWE-bench Multilingual | 89.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| SWE-bench Multimodal | 54.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| SWE-bench Pro (Anthropic internal subset) | 88.6% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 1 Oct 2026 | — | Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $0.54 per solved task |
| SWE-bench Pro (Anthropic internal subset) | 92.3% | Highadaptive thinking, effort=high (default) | Maker's own figureVendor-reportedAnthropic ↗ | 1 Oct 2026 | — | Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $1.19 per solved task |
| SWE-Bench Pro (public, v1) | 81.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| SWE-Bench Pro V2 (full) | 99.1% | HighFable 5.1 (Claude Code) high | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 2 (Scale rank accounts for CI); ±0.5; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-Bench Pro V2 (hard) | 92.2% | HighFable 5.1 (Claude Code) high | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 2 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| Terminal-Bench 2.1 | 85.0% | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-low) |
| Terminal-Bench 2.1 | 88.0% | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-medium) |
| Terminal-Bench 2.1 | 89.9% | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-high) |
| Terminal-Bench 2.1 | 85.0% | Higheffort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±0.991 stderr; $3.075/test |
| Terminal-Bench 2.1 | 91.0% | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-xhigh) |
| Terminal-Bench 2.1 | 91.4% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1) |
| Terminal-Bench 4.0 | 43.3% | Low | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.61 95% CI; 330 trials; total run cost $2359; model release 2026-09-01 |
| Terminal-Bench 4.0 | 40.4% | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-low) |
| Terminal-Bench 4.0 | 40.2% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; production safeguards on; vendor-reported cost/task $5.7 |
| Terminal-Bench 4.0 | 53.9% | Medium | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.39 95% CI; 330 trials; total run cost $2833; model release 2026-09-01 |
| Terminal-Bench 4.0 | 44.9% | MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-medium) |
| Terminal-Bench 4.0 | 43.4% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; production safeguards on; vendor-reported cost/task $7.8 |
| Terminal-Bench 4.0 | 54.5% | High | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.44 95% CI; 330 trials; total run cost $3985; model release 2026-09-01 |
| Terminal-Bench 4.0 | 52.0% | HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-high) |
| Terminal-Bench 4.0 | 49.4% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; production safeguards on; vendor-reported cost/task $10.5 |
| Terminal-Bench 4.0 | 57.9% | Extra high | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.36 95% CI; 330 trials; total run cost $4872; model release 2026-09-01 |
| Terminal-Bench 4.0 | 56.1% | Extra highClaude Fable 5.1 XHigh + SWE-2 Medium | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Devin Fusion CLI | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 55.1% | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1-xhigh) |
| Terminal-Bench 4.0 | 51.3% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; production safeguards on; vendor-reported cost/task $15.8 |
| Terminal-Bench 4.0 | 58.1% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±3.312 stderr; $17.175/test |
| Terminal-Bench 4.0 | 57.9% | Max | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.76 95% CI; 330 trials; total run cost $6244; model release 2026-09-01 |
| Terminal-Bench 4.0 | 57.6% | MaxFable 5.1 (max) (with fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 52.0% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-fable-5-1) |
| Terminal-Bench 4.0 | 55.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; production safeguards on; vendor-reported cost/task $19.5 |
| Terminal-Bench-Science 0.1 | 26.3% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Claude Code (--bare) | Terminal-Bench-Science 0.1 (70 tasks); SE ±3.5-4.5; vendor-reported cost/task $11.1 |
| Terminal-Bench-Science 0.1 | 35.7% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Claude Code (--bare) | Terminal-Bench-Science 0.1 (70 tasks); SE ±3.5-4.5; vendor-reported cost/task $14.9 |
| Terminal-Bench-Science 0.1 | 40.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Claude Code (--bare) | Terminal-Bench-Science 0.1 (70 tasks); SE ±3.5-4.5; vendor-reported cost/task $20.3 |
| Terminal-Bench-Science 0.1 | 49.5% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Claude Code (--bare) | Terminal-Bench-Science 0.1 (70 tasks); SE ±3.5-4.5; vendor-reported cost/task $31.8 |
| Terminal-Bench-Science 0.1 | 52.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Claude Code (--bare) | Terminal-Bench-Science 0.1 (70 tasks); SE ±3.5-4.5; vendor-reported cost/task $37.9 |
| Toolathlon | 77.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | internal harness, Pass@1 |
| Vals Code Migration | 54.6% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.812 stderr; $70.973/test |
| Vals Index | 65.8% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.115 stderr; $28.712/test |
| Vals Legal Research Bench | 55.3% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-fable-5-1; rank 3/72; ±3.456 stderr; $23.055789/test |
| Vals Public Benefits Bench v1.1 | 74.9% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-fable-5-1; rank 2/45; ±1.128 stderr; $7.481573/test |
| Vals SRE Bench | 22.9% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±2.601 stderr; $32.232/test |
| Vals TaxEval v2 | 76.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id anthropic/claude-fable-5-1; rank 9/145; ±0.83 stderr; $0.346549/test |
| Vals Vibe Code Bench | 90.3% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±1.566 stderr; $33.368/test |
| WANDR | 63.3% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $23.25 |
| WANDR | 64.5% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $27.41 |
| WANDR | 66.7% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $33.62 |
| WANDR | 67.7% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $42.81 |
| WANDR | 68.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $49.02 |