Skip to content
Bencher

Models · Anthropic · Out sinceReleased 1 Sep 2026

Claude Fable 5.1#4 for writing.#4 for writing, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)ProprietaryOur pick: Best for long, multi-step tasksPick: Best for long-running agents
Where you can use it:In apps:Claude (Team plan) · Fable 5.1

Claude Fable 5.1 is made by Anthropic. Among the models we track it ranks #4 for writing, #5 for research and analysis, #5 for coding. It's expensive to use. When an AI has to work on its own for a long time — using a computer, browsing, or handling a whole task end-to-end — this one is the most reliable.

API id claude-fable-5-1. Same model as Claude Mythos 5.1 with stricter safeguards (fallback to Opus models on cyber/bio). Default effort high (high in Claude Code, medium in Cowork/claude.ai). Cache reads $0.25/M (75% cheaper than Fable 5).

Writing & creativity
74.9 / 100 · #4
Research & analysis
80.1 / 100 · #5
Coding
81.8 / 100 · #5
Price
Expensive$10 / $50
Price per 1M (blended)Blended / 1M
$20
MemoryContext
1M
Longest answerMax output
128K
Test resultsResults
205 (135 independent135 indep.)
Out sinceReleased
1 Sep 2026
Made byVendor
Anthropic
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra highMax

Highlighted: where it did best for writing (Maximum thinking).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-fable-5.1.

Route it as anthropic/claude-fable-5.1 at $10 in / $50 out per 1M tokens, 1M context. Listed since 1 Sep 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstool_choicetoolsverbosity

Thinking level

Reasoning effort

How long should Claude Fable 5.1Where Claude Fable 5.1think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Artificial Analysis Coding Agent Index

Best at MaxBest at Max

Terminal-Bench 4.0

Best at Extra highBest at Extra high

CursorBench

Best at MaxBest at Max

DeepSWE

Best at MaxBest at Max

FrontierCode

Best at Medium, worse abovePeaks at Medium

CursorBench 3.2.0

Best at MaxBest at Max

SWE Atlas - Codebase QnA

Best at Extra high, worse abovePeaks at Extra high

Terminal-Bench 2.1

Best at MaxBest at Max

SciCode

Best at MaxBest at Max

SWE-bench Pro (Anthropic internal subset)

Best at HighBest at High

Artificial Analysis Intelligence Index

Best at MaxBest at Max

Terminal-Bench-Science 0.1

Best at MaxBest at Max

AA-Briefcase v1.1

Best at MaxBest at Max

AA-LCR

Best at MaxBest at Max

AA-Omniscience Accuracy

Best at MaxBest at Max

AA-Omniscience Hallucination Rate

Best at Low, worse abovePeaks at Low

Lower is better on this test

AA-Omniscience Index

Best at MaxBest at Max

ARC-AGI-2

Best at Extra highBest at Extra high

GDPval-AA v2.1

Best at MaxBest at Max

GPQA Diamond

Best at MaxBest at Max

Harvey LAB-AA

Best at Extra high, worse abovePeaks at Extra high

Humanity's Last Exam

Best at MaxBest at Max

Humanity's Last Exam (with tools)

Best at Extra high, worse abovePeaks at Extra high

WANDR

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent57.5%MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
AA-Briefcase1694Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—
AA-Briefcase v1.11669Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Briefcase v1.11678MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Briefcase v1.11678Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
AA-LCR82.3%LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR84.7%MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR83.7%HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR83.0%Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR85.3%MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy60.2%LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy63.1%MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy64.9%HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy66.2%Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy67.2%MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate65.6%LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate69.1%MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate68.8%HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate70.5%Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate72.6%MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index34.1LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index37.6MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index40.8HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index42.4Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index43.5MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
ARC-AGI-197.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—
ARC-AGI-286.2%MediumClaude Fable 5.1 (Medium)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $1.22/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-288.8%HighClaude Fable 5.1 (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $1.67/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-290.0%Extra highClaude Fable 5.1 (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $3.12/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-290.0%MaxClaude Fable 5.1 (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $4.49/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-290.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—
Artificial Analysis Coding Agent Index61.7Extra highClaude Fable 5.1 XHigh + SWE-2 MediumIndependent testIndependentArtificial Analysis ↗1 Oct 2026Devin Fusion CLIagent Devin Fusion CLI; components: DeepSWE v1.1 63.1, SWE-Atlas-QnA 65.9, Terminal-Bench v4 56.1; avg cost $7.90/task; avg wall time 36 min/task; Devin Fusion = Fable/GPT planner + Cognition SWE-2 (medium) sidekick
Artificial Analysis Coding Agent Index62.2MaxFable 5.1 (max) (with fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude Codeagent Claude Code; components: DeepSWE v1.1 64.3, SWE-Atlas-QnA 64.8, Terminal-Bench v4 57.6; avg cost $12.39/task; avg wall time 35 min/task; with fallback model
Artificial Analysis Intelligence Index46.8LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-fable-5-1-low; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $2.37/task
Artificial Analysis Intelligence Index48.9MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-fable-5-1-medium; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $2.98/task
Artificial Analysis Intelligence Index51.2HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-fable-5-1-high; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $3.91/task
Artificial Analysis Intelligence Index53.2Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-fable-5-1-xhigh; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $5.98/task
Artificial Analysis Intelligence Index53.4MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-fable-5-1; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $7.63/task
Artificial Analysis output speed47 tok/sLowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 4.5s; list price $10/50 per 1M in/out
Artificial Analysis output speed49 tok/sMediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 8.1s; list price $10/50 per 1M in/out
Artificial Analysis output speed51 tok/sHighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 29.7s; list price $10/50 per 1M in/out
Artificial Analysis output speed59 tok/sExtra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 102.1s; list price $10/50 per 1M in/out
Artificial Analysis output speed68 tok/sMaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 262.3s; list price $10/50 per 1M in/out
ArXivMath (no tools)82.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—Aug 2026
ArXivMath (with tools)92.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—Aug 2026
AutomationBench31.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps)
Chartography (with tools)88.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
CursorBench45.1%LowFable 5.1 LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 15; $5.44/task; 34,795 tokens/task; 51 steps/task
CursorBench45.1%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $5.44
CursorBench46.8%MediumFable 5.1 MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 11; $7.05/task; 45,411 tokens/task; 63 steps/task
CursorBench46.8%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $7.05
CursorBench49.2%HighFable 5.1 HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 9; $9.08/task; 58,438 tokens/task; 77 steps/task
CursorBench49.2%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $9.08
CursorBench51.6%Extra highFable 5.1 Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 8; $13.01/task; 87,294 tokens/task; 101 steps/task
CursorBench51.6%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $13.01
CursorBench51.8%MaxFable 5.1 MaxIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 7; $17.28/task; 117,236 tokens/task; 128 steps/task
CursorBench51.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $17.28
CursorBench 3.2.066.2%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $2.9
CursorBench 3.2.068.0%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $3.53
CursorBench 3.2.069.4%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $4.8
CursorBench 3.2.072.8%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $6.96
CursorBench 3.2.073.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $9.64
DeepSWE63.1%Extra highClaude Fable 5.1 XHigh + SWE-2 MediumIndependent testIndependentArtificial Analysis ↗1 Oct 2026Devin Fusion CLIAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE64.3%MaxFable 5.1 (max) (with fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE67.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—
Design Arena (all categories)1342Defaultclaude-fable-5-1Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±3.8 SE; 9335 battles; win rate 59.3%
Design Arena (fullstack)1324Defaultclaude-fable-5-1Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±15.9 SE; 509 battles; win rate 57.2%
Epoch Capabilities Index164.8DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 161.8-169.1; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)2162Default*claude-fable-5-1Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 2; rubric score 16.95/20; slop 8.16; avg length 5841 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*)
FrontierCode49.8%Loweffort=lowIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $2.38; competitor number as plotted by OpenAI
FrontierCode52.8%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Fable 5.1 adds unrequested out-of-scope edits at higher effort; vendor-reported cost/task $2.47
FrontierCode50.9%Mediumclaude-fable-5-1_mediumIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026Claude Coderead from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5
FrontierCode50.9%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Fable 5.1 adds unrequested out-of-scope edits at higher effort; vendor-reported cost/task $3.28
FrontierCode50.3%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Fable 5.1 adds unrequested out-of-scope edits at higher effort; vendor-reported cost/task $5.27
FrontierCode48.7%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Fable 5.1 adds unrequested out-of-scope edits at higher effort; vendor-reported cost/task $9.27
FrontierCode50.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Fable 5.1 adds unrequested out-of-scope edits at higher effort; vendor-reported cost/task $12.82
FrontierCode v1.1 (Extended)63.6%Mediumadaptive thinking, effort=medium (best)Maker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude CodeFable 5.1 peaks at medium on FrontierCode; adds unrequested out-of-scope edits at higher effort
FrontierMath (Tiers 1-3)90.2%Maxclaude-fable-5-1_maxIndependent testIndependentEpoch AI ↗1 Sep 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 487.8%Maxclaude-fable-5-1_maxIndependent testIndependentEpoch AI ↗1 Sep 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 5.2pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierSWE v256.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Proximal harnessreported as 0.57 on a 0-1 scale (Opus 5.5 card later lists 56.3%)
GDPval-AA v21853Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—GDPval-AA v2 Elo
GDPval-AA v2.11450Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $1.41
GDPval-AA v2.11536Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $2.17
GDPval-AA v2.11617Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $3.43
GDPval-AA v2.11721Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1695.77-1746.08
GDPval-AA v2.11721Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $7.09
GDPval-AA v2.11735MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1710.09-1759.25
GDPval-AA v2.11735Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $9.59
GPQA Diamond88.1%LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-low)
GPQA Diamond88.6%MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-medium)
GPQA Diamond90.6%HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-high)
GPQA Diamond93.4%Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-xhigh)
GPQA Diamond93.7%MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1)
GPQA Diamond93.4%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±1.863 stderr; $0.205/test
GSO88.2%Extra highreasoning_effort=xhighIndependent testIndependentGSO ↗27 Sep 2026OpenHandsOpt@1; hack-controlled score 87.25; data https://gso-bench.github.io/assets/leaderboard.json
Harvey LAB-AA93.0%HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
Harvey LAB-AA93.3%Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
Harvey LAB-AA93.0%MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
HealthBench Professional62.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—
HLE Diamond51.3%Defaultclaude-fable-5-1Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±3.1; entry added 2026-04-10
Humanity's Last Exam48.9%LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-low)
Humanity's Last Exam53.2%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam without tools; vendor-reported cost/task $0.3
Humanity's Last Exam53.8%MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-medium)
Humanity's Last Exam55.9%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam without tools; vendor-reported cost/task $0.46
Humanity's Last Exam55.9%HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-high)
Humanity's Last Exam58.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam without tools; vendor-reported cost/task $0.75
Humanity's Last Exam58.7%Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-xhigh)
Humanity's Last Exam46.5%Extra highFable 5.1 (xhigh)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±2; entry added 2026-09-03
Humanity's Last Exam60.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam without tools; vendor-reported cost/task $1.53
Humanity's Last Exam59.1%MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1)
Humanity's Last Exam60.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam without tools; vendor-reported cost/task $2.23
Humanity's Last Exam (with tools)60.0%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam with tools (web search/fetch, code execution); vendor-reported cost/task $0.52
Humanity's Last Exam (with tools)63.0%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam with tools (web search/fetch, code execution); vendor-reported cost/task $0.67
Humanity's Last Exam (with tools)64.8%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam with tools (web search/fetch, code execution); vendor-reported cost/task $1.05
Humanity's Last Exam (with tools)65.1%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam with tools (web search/fetch, code execution); vendor-reported cost/task $2.28
Humanity's Last Exam (with tools)65.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam with tools (web search/fetch, code execution); vendor-reported cost/task $3.2
IOI (Vals)90.8%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.647 stderr; $11.197/test
LegalBench (Vals)88.5%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-fable-5-1; rank 2/149; ±0.42 stderr; $0.051606/test
LiveCodeBench90.5%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±0.863 stderr; $0.261/test
LLM Creative Story-Writing Benchmark (Lech Mazur)3.8HighClaude Fable 5.1 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 1/56; Thurstone comparison score (centered at 0); est. win chance 93%; 95% bootstrap 3.702 to 3.896
LMArena Code Arena (WebDev)1751Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 4 (CI rank 3-4); 95% CI 1741-1760; 6137 votes
LMArena Text - Coding category1533Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 17 (CI rank 2-48); 95% CI 1520-1545; 2254 votes
LMArena Text - Creative Writing1482Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 10 (rank range 4-26); 95% CI 1468.8-1494.5; 2506 votes; style-controlled
LMArena Text - Expert1519Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 16 (rank range 3-62); 95% CI 1499.9-1538.6; 965 votes; style-controlled
LMArena Text - Hard Prompts1522Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 7 (rank range 2-19); 95% CI 1513.6-1529.4; 6630 votes; style-controlled
LMArena Text - Instruction Following1498Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 7 (rank range 3-20); 95% CI 1486.9-1508.5; 3380 votes; style-controlled
LMArena Text - Longer Query1511Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 7 (rank range 2-21); 95% CI 1502.0-1520.8; 4705 votes; style-controlled
LMArena Text - Multi-Turn1487Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 25 (rank range 7-74); 95% CI 1471.1-1502.7; 1474 votes; style-controlled
LMArena Text - Non-English1495Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 3 (rank range 1-13); 95% CI 1487.6-1503.2; 6689 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1487Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 15 (rank range 3-48); 95% CI 1473.7-1499.8; 2102 votes; style-controlled
LMArena Text - Occupational: Legal & Government1494Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 19 (rank range 1-74); 95% CI 1474.2-1514.4; 918 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1493Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 7 (rank range 2-19); 95% CI 1481.5-1504.0; 3159 votes; style-controlled
LMArena Text (overall)1501Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 6 (CI rank 2-14); 95% CI 1495-1508; 11241 votes
MCP Atlas87.2%DefaultFable 5.1Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 1 (Scale rank accounts for CI); ±2.05; entry added 2026-09-04
OfficeQA Pro69.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
OSWorld 2.077.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—partial score; August 2026 task release; safeguard interventions scored 0
OSWorld 2.0 (strict pass rate)41.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—strict pass rate; August 2026 task release
OSWorld 2.1 (partial score)80.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—partial score; re-evaluated by Anthropic under Opus 5.5 conditions
OSWorld 2.1 (strict pass rate)42.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
PRBench Finance (Scale)50.8%DefaultFable 5.1Independent testIndependentScale AI (SEAL) ↗3 Sep 2026—Scale rank 7; ±0.22 CI
PRBench Legal (Scale)51.6%DefaultFable 5.1Independent testIndependentScale AI (SEAL) ↗3 Sep 2026—Scale rank 5; ±0.24 CI
ProgramBench (fully resolved)7.0%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.809 stderr; $58.517/test; strict fully-resolved rate
ProgramBench (fully resolved)87.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026mini-swe-agent
SciCode56.7%LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-low)
SciCode56.4%MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-medium)
SciCode58.7%HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-high)
SciCode60.9%Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-xhigh)
SciCode63.1%MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1)
SimpleBench86.6%DefaultClaude Fable 5.1Independent testIndependentSimpleBench ↗3 Sep 2026—AVG@5, temp 0.7; rank 2nd; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified70.8%Maxclaude-fable-5-1_maxIndependent testIndependentEpoch AI ↗1 Sep 2026—Epoch-run (no tools); ±1.44 stderr
SkillsBench61.5%Maxeffort=maxIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.76 stderr; $4.517/test
SWE Atlas - Codebase QnA65.9%Extra highClaude Fable 5.1 XHigh + SWE-2 MediumIndependent testIndependentArtificial Analysis ↗1 Oct 2026Devin Fusion CLIAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA60.0%Extra highFable-5.1 (Claude Code) xHigh*Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±4.85; entry added 2026-03-02; * = Scale footnote (see page)
SWE Atlas - Codebase QnA64.8%MaxFable 5.1 (max) (with fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Refactoring56.7%Extra highFable-5.1 (Claude Code) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±6.52; entry added 2026-09-02
SWE Atlas - Test Writing67.0%Extra highFable-5.1 (Claude Code) xHigh*Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±5.33; entry added 2026-09-02; * = Scale footnote (see page)
SWE-bench Multilingual89.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
SWE-bench Multimodal54.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
SWE-bench Pro (Anthropic internal subset)88.6%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗1 Oct 2026—Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $0.54 per solved task
SWE-bench Pro (Anthropic internal subset)92.3%Highadaptive thinking, effort=high (default)Maker's own figureVendor-reportedAnthropic ↗1 Oct 2026—Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $1.19 per solved task
SWE-Bench Pro (public, v1)81.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
SWE-Bench Pro V2 (full)99.1%HighFable 5.1 (Claude Code) highIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 2 (Scale rank accounts for CI); ±0.5; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-Bench Pro V2 (hard)92.2%HighFable 5.1 (Claude Code) highIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 2 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
Terminal-Bench 2.185.0%LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-low)
Terminal-Bench 2.188.0%MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-medium)
Terminal-Bench 2.189.9%HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-high)
Terminal-Bench 2.185.0%Higheffort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.991 stderr; $3.075/test
Terminal-Bench 2.191.0%Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-xhigh)
Terminal-Bench 2.191.4%MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1)
Terminal-Bench 4.043.3%LowIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.61 95% CI; 330 trials; total run cost $2359; model release 2026-09-01
Terminal-Bench 4.040.4%LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-low)
Terminal-Bench 4.040.2%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; production safeguards on; vendor-reported cost/task $5.7
Terminal-Bench 4.053.9%MediumIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.39 95% CI; 330 trials; total run cost $2833; model release 2026-09-01
Terminal-Bench 4.044.9%MediumClaude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-medium)
Terminal-Bench 4.043.4%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; production safeguards on; vendor-reported cost/task $7.8
Terminal-Bench 4.054.5%HighIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.44 95% CI; 330 trials; total run cost $3985; model release 2026-09-01
Terminal-Bench 4.052.0%HighClaude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-high)
Terminal-Bench 4.049.4%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; production safeguards on; vendor-reported cost/task $10.5
Terminal-Bench 4.057.9%Extra highIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.36 95% CI; 330 trials; total run cost $4872; model release 2026-09-01
Terminal-Bench 4.056.1%Extra highClaude Fable 5.1 XHigh + SWE-2 MediumIndependent testIndependentArtificial Analysis ↗1 Oct 2026Devin Fusion CLIAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.055.1%Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1-xhigh)
Terminal-Bench 4.051.3%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; production safeguards on; vendor-reported cost/task $15.8
Terminal-Bench 4.058.1%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±3.312 stderr; $17.175/test
Terminal-Bench 4.057.9%MaxIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.76 95% CI; 330 trials; total run cost $6244; model release 2026-09-01
Terminal-Bench 4.057.6%MaxFable 5.1 (max) (with fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.052.0%MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5-1)
Terminal-Bench 4.055.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; production safeguards on; vendor-reported cost/task $19.5
Terminal-Bench-Science 0.126.3%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); SE ±3.5-4.5; vendor-reported cost/task $11.1
Terminal-Bench-Science 0.135.7%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); SE ±3.5-4.5; vendor-reported cost/task $14.9
Terminal-Bench-Science 0.140.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); SE ±3.5-4.5; vendor-reported cost/task $20.3
Terminal-Bench-Science 0.149.5%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); SE ±3.5-4.5; vendor-reported cost/task $31.8
Terminal-Bench-Science 0.152.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); SE ±3.5-4.5; vendor-reported cost/task $37.9
Toolathlon77.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—internal harness, Pass@1
Vals Code Migration54.6%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.812 stderr; $70.973/test
Vals Index65.8%Maxeffort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.115 stderr; $28.712/test
Vals Legal Research Bench55.3%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-fable-5-1; rank 3/72; ±3.456 stderr; $23.055789/test
Vals Public Benefits Bench v1.174.9%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-fable-5-1; rank 2/45; ±1.128 stderr; $7.481573/test
Vals SRE Bench22.9%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±2.601 stderr; $32.232/test
Vals TaxEval v276.0%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id anthropic/claude-fable-5-1; rank 9/145; ±0.83 stderr; $0.346549/test
Vals Vibe Code Bench90.3%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±1.566 stderr; $33.368/test
WANDR63.3%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $23.25
WANDR64.5%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $27.41
WANDR66.7%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $33.62
WANDR67.7%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $42.81
WANDR68.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $49.02