Skip to content
Bencher

Models · Anthropic · Out sinceReleased 9 Jun 2026

Claude Fable 5#3 for writing.#3 for writing, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary
Where you can use it:In apps:Claude (Team plan) · Fable 5

Claude Fable 5 is made by Anthropic. Among the models we track it ranks #3 for writing, #7 for research and analysis, #8 for coding. It's expensive to use.

API id claude-fable-5. Same model as Claude Mythos 5 with safeguards (fallback to Opus 4.8). Default effort high; no per-message effort.

Writing & creativity
78.2 / 100 · #3
Research & analysis
79.8 / 100 · #7
Coding
71.8 / 100 · #8
Price
Expensive$10 / $50
Price per 1M (blended)Blended / 1M
$20
MemoryContext
1M
Longest answerMax output
128K
Test resultsResults
124 (79 independent79 indep.)
Out sinceReleased
9 Jun 2026
Made byVendor
Anthropic
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra highMax

Highlighted: where it did best for writing (High thinking).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-fable-5.

Route it as anthropic/claude-fable-5 at $10 in / $50 out per 1M tokens, 1M context. Listed since 9 Jun 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstool_choicetoolsverbosity

Thinking level

Reasoning effort

How long should Claude Fable 5Where Claude Fable 5think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Artificial Analysis Coding Agent Index

Best at MaxBest at Max

DeepSWE

Best at Extra high, worse abovePeaks at Extra high

CursorBench 3.2.0

Best at MaxBest at Max

Terminal-Bench 2.1

Best at MaxBest at Max

Terminal-Bench-Science 0.1

Best at High, worse abovePeaks at High

ARC-AGI-2

Best at MaxBest at Max

Humanity's Last Exam

Best at Extra high, worse abovePeaks at Extra high

Humanity's Last Exam (with tools)

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent48.8%MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
AA-Briefcase1574Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—
Agents' Last Exam48.7%Extra higheffort=xhighIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $28.5542; competitor number as plotted by OpenAI ("Claude Fable 5 w/ Opus 4.8 fallback")
Agents' Last Exam41.3%Defaultadaptive (as labeled)Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $15.2285; competitor number as plotted by OpenAI ("Claude Fable 5 w/ Opus 4.8 fallback")
ARC-AGI-198.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—
ARC-AGI-287.5%HighClaude Fable 5 (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $2.44/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-288.3%Extra highClaude Fable 5 (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $3.00/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-289.2%MaxClaude Fable 5 (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $5.45/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-289.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—
Artificial Analysis Coding Agent Index58.6Loweffort=lowIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.17; competitor number as plotted by OpenAI
Artificial Analysis Coding Agent Index63.2Mediumeffort=mediumIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $4.74; competitor number as plotted by OpenAI
Artificial Analysis Coding Agent Index65.1Higheffort=highIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $5.97; competitor number as plotted by OpenAI
Artificial Analysis Coding Agent Index66Extra higheffort=xhighIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $8.53; competitor number as plotted by OpenAI
Artificial Analysis Coding Agent Index67.2Maxeffort=maxIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $11.69; competitor number as plotted by OpenAI
Artificial Analysis Intelligence Index49.6MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-fable-5; list price $10/50 per 1M in/out; cost to run AA Intelligence Index $8.75/task (AA marks this variant deprecated)
Artificial Analysis output speed63 tok/sMaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 96.0s; list price $10/50 per 1M in/out (AA marks this variant deprecated)
AutomationBench17.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—read from launch-post benchmark table image (values printed in table)
Blueprint-Bench 238.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026—read from launch-post benchmark table image (values printed in table)
BrowseComp87.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
CursorBench (pre-3.2, legacy)72.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026Cursor agentCursorBench version as of June 2026; leads at every effort from medium up
CursorBench 3.2.062.1%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $4.46
CursorBench 3.2.065.2%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $6.8
CursorBench 3.2.066.5%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $8.77
CursorBench 3.2.068.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $11.73
CursorBench 3.2.070.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0; vendor-reported cost/task $17.32
DeepSWE59.6%Lowclaude-fable-5_lowIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 81.4%; ±2.8; 4 runs; $3.76/task
DeepSWE59.6%Loweffort=lowIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $3.7579; competitor number as plotted by OpenAI
DeepSWE65.4%Mediumclaude-fable-5_mediumIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 83.2%; ±4.4; 4 runs; $6.09/task
DeepSWE65.4%Mediumeffort=mediumIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $6.0882; competitor number as plotted by OpenAI
DeepSWE68.6%Highclaude-fable-5_highIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 86.7%; ±1.1; 4 runs; $9.18/task
DeepSWE68.6%Higheffort=highIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $9.1776; competitor number as plotted by OpenAI
DeepSWE69.9%Extra higheffort=xhighIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $13.4145; competitor number as plotted by OpenAI
DeepSWE69.9%Extra highclaude-fable-5_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 88.5%; ±3.2; 4 runs; $13.41/task
DeepSWE69.7%Maxclaude-fable-5_maxIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 84.1%; ±4.0; 4 runs; $21.63/task
DeepSWE69.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
Design Arena (all categories)1317Defaultclaude-fable-5Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.7 SE; 20185 battles; win rate 57.7%
Design Arena (fullstack)1278Defaultclaude-fable-5Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±8.2 SE; 2118 battles; win rate 58.2%
Epoch Capabilities Index162.2DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 159.6-166.1; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1943Defaultclaude-fable-5Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 11; rubric score 16.81/20; slop 10.28; avg length 5887 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
Frontier-Bench v0.133.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
FrontierCode53.5%Extra highadaptive thinking, effort=xhigh (best)Maker's own figureVendor-reportedAnthropic ↗24 Jul 2026Claude CodeFrontierCode v1.1 Main, each model at best effort (Opus 5.5 system card: Fable 5 53.5% at xhigh)
FrontierCode46.3%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026Claude CodeFrontierCode Main at Fable 5 launch, metric = score (pass rate 48.8%); earlier FrontierCode version than v1.1 numbers
FrontierCode53.5%Defaultclaude-fable-5_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026Claude Coderead from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5
FrontierCode (Diamond)29.3%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026Claude CodeFrontierCode Diamond subset, score (pass rate 30.2%); read from launch-post benchmark table image (values printed in table)
FrontierMath (Tiers 1-3)87.0%Maxclaude-fable-5_maxIndependent testIndependentEpoch AI ↗9 Jun 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.0pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 490.2%Maxclaude-fable-5_maxIndependent testIndependentEpoch AI ↗9 Jun 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 4.6pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierSWE v248.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Proximal harnessreported as 0.48
GDP.pdf29.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026—no tools; read from launch-post benchmark table image (values printed in table)
GDPval-AA (v1)1932Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026—GDPval-AA (v1) Elo; higher of Mythos 5/Fable 5; read from launch-post benchmark table image (values printed in table)
GDPval-AA v21723Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—read from launch-post benchmark table image (values printed in table)
GPQA Diamond93.2%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±1.939 stderr; $0.450/test
GPQA Diamond92.6%MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5) (AA marks this variant deprecated)
GSO78.4%Extra highreasoning_effort=xhighIndependent testIndependentGSO ↗27 Sep 2026OpenHandsOpt@1; hack-controlled score 76.47; data https://gso-bench.github.io/assets/leaderboard.json
Harvey LAB-AA93.6%MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
Humanity's Last Exam50.6%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam without tools; vendor-reported cost/task $0.17
Humanity's Last Exam55.9%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam without tools; vendor-reported cost/task $0.4
Humanity's Last Exam56.9%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam without tools; vendor-reported cost/task $0.62
Humanity's Last Exam57.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam without tools; vendor-reported cost/task $0.91
Humanity's Last Exam55.5%MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5) (AA marks this variant deprecated)
Humanity's Last Exam57.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam without tools; vendor-reported cost/task $1.7
Humanity's Last Exam (with tools)59.6%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam with tools; vendor-reported cost/task $0.61
Humanity's Last Exam (with tools)61.4%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam with tools; vendor-reported cost/task $1.01
Humanity's Last Exam (with tools)63.2%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam with tools; vendor-reported cost/task $1.42
Humanity's Last Exam (with tools)63.5%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam with tools; vendor-reported cost/task $1.92
Humanity's Last Exam (with tools)63.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—Humanity's Last Exam with tools; vendor-reported cost/task $3.44
LegalBench (Vals)88.6%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-fable-5; rank 1/149; ±0.331 stderr; $0.030639/test
LiveCodeBench89.8%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±0.892 stderr; $0.428/test
LLM Creative Story-Writing Benchmark (Lech Mazur)2.9HighClaude Fable 5 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 7/56; Thurstone comparison score (centered at 0); est. win chance 86%; 95% bootstrap 2.803 to 2.946; incomplete story set (see README coverage note)
LMArena Code Arena (WebDev)1626Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 17 (CI rank 14-24); 95% CI 1619-1632; 13952 votes
LMArena Search Arena1230Highclaude-fable-5-highIndependent testIndependentLMArena ↗24 Aug 2026—rank 5 (rank range 3-7); 95% CI 1222.3-1238.1; 41795 votes
LMArena Text - Coding category1552Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 2 (CI rank 1-13); 95% CI 1545-1559; 9946 votes
LMArena Text - Creative Writing1503Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 3 (rank range 1-6); 95% CI 1495.3-1510.7; 8072 votes; style-controlled
LMArena Text - Expert1549Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 1 (rank range 1-13); 95% CI 1539.3-1558.7; 4217 votes; style-controlled
LMArena Text - Hard Prompts1531Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 4 (rank range 2-7); 95% CI 1525.5-1535.7; 24800 votes; style-controlled
LMArena Text - Instruction Following1510Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 4 (rank range 1-7); 95% CI 1503.5-1515.9; 13415 votes; style-controlled
LMArena Text - Longer Query1523Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 4 (rank range 2-7); 95% CI 1517.2-1528.8; 17723 votes; style-controlled
LMArena Text - Multi-Turn1518Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 4 (rank range 2-10); 95% CI 1509.4-1525.9; 6231 votes; style-controlled
LMArena Text - Non-English1494Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 4 (rank range 2-12); 95% CI 1488.7-1499.2; 21613 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1502Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 5 (rank range 1-15); 95% CI 1494.4-1509.5; 7646 votes; style-controlled
LMArena Text - Occupational: Legal & Government1512Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 5 (rank range 1-28); 95% CI 1500.4-1522.9; 3185 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1509Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 3 (rank range 1-7); 95% CI 1502.0-1515.6; 10571 votes; style-controlled
LMArena Text (overall)1505Highclaude-fable-5-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 3 (CI rank 2-9); 95% CI 1500-1509; 37900 votes
MCP Atlas83.3%DefaultClaude Fable 5Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±2.25; entry added 2026-06-09
OSWorld 2.072.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—partial score; August 2026 task release; safeguard interventions scored 0
OSWorld 2.0 (strict pass rate)36.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—strict pass rate; August 2026 task release
OSWorld-Verified85.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026—read from launch-post benchmark table image (values printed in table)
PRBench Finance (Scale)53.9%Defaultclaude-fable 5Independent testIndependentScale AI (SEAL) ↗28 Jul 2026—Scale rank 3; ±0.14 CI
PRBench Legal (Scale)52.6%Defaultclaude-fable 5Independent testIndependentScale AI (SEAL) ↗28 Jul 2026—Scale rank 3; ±0.54 CI
ProgramBench (fully resolved)2.0%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.992 stderr; $75.676/test; strict fully-resolved rate
SciCode61.0%MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5) (AA marks this variant deprecated)
SimpleBench81.9%DefaultClaude FableIndependent testIndependentSimpleBench ↗10 Jun 2026—AVG@5, temp 0.7; rank 6th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified70.7%Extra highclaude-fable-5_xhighIndependent testIndependentEpoch AI ↗10 Aug 2026—Epoch-run (no tools); ±1.44 stderr
SWE Atlas - Codebase QnA39.0%Extra highFable-5 (Claude Code) xHigh*Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 5 (Scale rank accounts for CI); ±5; entry added 2026-03-02; * = Scale footnote (see page)
SWE Atlas - Refactoring54.8%Extra highFable-5 (Claude Code) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±6.76; entry added 2026-06-11
SWE Atlas - Test Writing55.6%Extra highFable-5 (Claude Code) xHigh*Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±5.8; entry added 2026-06-11; * = Scale footnote (see page)
SWE-bench Multilingual86.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—
SWE-bench Multimodal54.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—
SWE-Bench Pro (public, v1)80.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—
SWE-bench Verified95.0%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±0.976 stderr; $2.047/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
SWE-bench Verified95.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026—
SWE-rebench64.5%HighFable 5 [high]Independent testIndependentSWE-rebench (Nebius) ↗1 Oct 2026SWE-rebench standard scaffoldtime window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.41; pass@5 78.4%; $4.40/problem
tau2-bench98.5%MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5) (AA marks this variant deprecated)
Terminal-Bench 2.184.3%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026mini-SWE-agent (GKE)20.9% of trials hit a safety refusal and fell back to Opus 4.8
Terminal-Bench 2.184.6%MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5) (AA marks this variant deprecated)
Terminal-Bench 2.180.5%Maxeffort=maxIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±1.35 stderr; $1.429/test
Terminal-Bench 2.183.8%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Claude CodeTB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.180.4%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Terminus 2TB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 3.034.1%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Claude CodeTB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 3.6B, run cost $6.5k
Terminal-Bench 4.044.5%MaxIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.85 95% CI; 330 trials; total run cost $7265; model release 2026-06-09
Terminal-Bench 4.042.4%MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-fable-5) (AA marks this variant deprecated)
Terminal-Bench 4.041.4%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.336 stderr; $30.340/test
Terminal-Bench 4.042.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)public leaderboard: 44.5%
Terminal-Bench-Science 0.112.3%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); vendor-reported cost/task $17.1
Terminal-Bench-Science 0.121.4%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); vendor-reported cost/task $25.0
Terminal-Bench-Science 0.125.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); vendor-reported cost/task $34.3
Terminal-Bench-Science 0.123.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); vendor-reported cost/task $36.0
Terminal-Bench-Science 0.124.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); vendor-reported cost/task $44.1
Vals Code Migration55.1%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.605 stderr; $112.098/test
Vals CorpFin v271.8%Maxeffort=maxIndependent testIndependentVals.ai ↗12 Aug 2026—vals id anthropic/claude-fable-5; rank 2/134; ±0.883 stderr; $1.737573/test
Vals Index61.4%Maxeffort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±0.999 stderr; $29.585/test
Vals Legal Research Bench49.5%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-fable-5; rank 6/72; ±3.475 stderr; $9.791426/test
Vals Public Benefits Bench v1.170.4%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-fable-5; rank 4/45; ±1.187 stderr; $4.585446/test
Vals TaxEval v276.9%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id anthropic/claude-fable-5; rank 5/145; ±0.821 stderr; $0.220133/test
Vals Vibe Code Bench90.3%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±2.097 stderr; $41.715/test