Skip to content
Bencher

Models · Anthropic · Out sinceReleased 24 Jul 2026

Claude Opus 5#2 for research and analysis.#2 for research and analysis, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary
Where you can use it:In apps:Claude (Team plan) · Opus 5

Claude Opus 5 is made by Anthropic. Among the models we track it ranks #2 for research and analysis, #6 for writing, #6 for coding. It's expensive to use.

API id claude-opus-5. Default effort high; thinking cannot be disabled at xhigh/max. Previous-generation Opus.

Writing & creativity
71.1 / 100 · #6
Research & analysis
84.7 / 100 · #2
Coding
80.5 / 100 · #6
Price
Expensive$5 / $25
Price per 1M (blended)Blended / 1M
$10
MemoryContext
1M
Longest answerMax output
128K
Test resultsResults
222 (153 independent153 indep.)
Out sinceReleased
24 Jul 2026
Made byVendor
Anthropic
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra highMax

Highlighted: where it did best for research and analysis (Maximum thinking).

Highlighted: dominant setting in its research and analysis composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-opus-5.

Route it as anthropic/claude-opus-5 at $5 in / $25 out per 1M tokens, 1M context. Listed since 24 Jul 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolsverbosity

Thinking level

Reasoning effort

How long should Claude Opus 5Where Claude Opus 5think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Artificial Analysis Coding Agent Index

Best at Extra high, worse abovePeaks at Extra high

Terminal-Bench 4.0

Best at MaxBest at Max

CursorBench

Best at MaxBest at Max

DeepSWE

Best at MaxBest at Max

FrontierCode

Best at MediumBest at Medium

Frontier-Bench v0.1

Best at Extra high, worse abovePeaks at Extra high

LMArena Code Arena (WebDev)

Best at MaxBest at Max

SWE Atlas - Codebase QnA

Best at Extra high, worse abovePeaks at Extra high

ProgramBench (fully resolved)

Best at Extra high, worse abovePeaks at Extra high

Terminal-Bench 3.0

Best at MaxBest at Max

Terminal-Bench 2.1

Best at MaxBest at Max

LMArena Text - Coding category

Best at High, worse abovePeaks at High

SciCode

Best at MaxBest at Max

Agents' Last Exam

Best at High, worse abovePeaks at High

Artificial Analysis Intelligence Index

Best at MaxBest at Max

AutomationBench

Best at MaxBest at Max

OSWorld 2.0

Best at MaxBest at Max

OSWorld 2.0 offline set

Best at MaxBest at Max

AA-Briefcase

Best at MaxBest at Max

ARC-AGI-2

Best at MaxBest at Max

GDPval-AA v2

Best at Extra high, worse abovePeaks at Extra high

GDPval-AA v2.1

Best at MaxBest at Max

GPQA Diamond

Best at High, worse abovePeaks at High

Humanity's Last Exam

Best at MaxBest at Max

LLM Creative Story-Writing Benchmark (Lech Mazur)

Best at Extra highBest at Extra high

LMArena Text - Creative Writing

Best at High, worse abovePeaks at High

LMArena Text - Expert

Best at High, worse abovePeaks at High

LMArena Text - Hard Prompts

Best at High, worse abovePeaks at High

LMArena Text - Instruction Following

Best at High, worse abovePeaks at High

LMArena Text - Longer Query

Best at High, worse abovePeaks at High

LMArena Text - Non-English

Best at High, worse abovePeaks at High

LMArena Text - Occupational: Legal & Government

Best at High, worse abovePeaks at High

LMArena Text - Occupational: Writing, Literature & Language

Best at High, worse abovePeaks at High

LMArena Text (overall)

Best at High, worse abovePeaks at High

WANDR

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent53.8%MaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
AA-Briefcase1606Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—AA-Briefcase (pre-v1.1), run by Artificial Analysis; xhigh/high use 15%/33% fewer output tokens than max
AA-Briefcase1693Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—AA-Briefcase (pre-v1.1), run by Artificial Analysis; xhigh/high use 15%/33% fewer output tokens than max
AA-Briefcase1720Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—AA-Briefcase (pre-v1.1), run by Artificial Analysis; xhigh/high use 15%/33% fewer output tokens than max
AA-Briefcase v1.11673MaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Briefcase v1.11673Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
Agents' Last Exam51.9%Loweffort=lowIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $4.0619; competitor number as plotted by OpenAI
Agents' Last Exam53.0%Mediumeffort=mediumIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $5.2907; competitor number as plotted by OpenAI
Agents' Last Exam55.9%Higheffort=highIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $7.2917; competitor number as plotted by OpenAI
Agents' Last Exam55.5%Extra higheffort=xhighIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $10.0226; competitor number as plotted by OpenAI
Agents' Last Exam52.7%Maxeffort=maxIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $9.7583; competitor number as plotted by OpenAI
ARC-AGI-197.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—
ARC-AGI-288.3%HighClaude Opus 5 (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $1.45/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-290.4%MaxClaude Opus 5 (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $2.06/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-290.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—semi-private set
ARC-AGI-330.2%HighClaude Opus 5 (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $20,657; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-330.2%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—ARC Prize Foundation verified, semi-private set (RHAE 30.16%); max-effort result not available at release
Artificial Analysis Coding Agent Index59.2No reasoningeffort=noneIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.53; competitor number as plotted by OpenAI
Artificial Analysis Coding Agent Index59.4Loweffort=lowIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $2.3; competitor number as plotted by OpenAI
Artificial Analysis Coding Agent Index64.1Mediumeffort=mediumIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.17; competitor number as plotted by OpenAI
Artificial Analysis Coding Agent Index65.6Higheffort=highIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.92; competitor number as plotted by OpenAI
Artificial Analysis Coding Agent Index68.1Extra higheffort=xhighIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $8.17; competitor number as plotted by OpenAI
Artificial Analysis Coding Agent Index67Maxeffort=maxIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗3 Sep 2026—Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $8.94; competitor number as plotted by OpenAI
Artificial Analysis Coding Agent Index59.7MaxOpus 5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude Codeagent Claude Code; components: DeepSWE v1.1 62.5, SWE-Atlas-QnA 62.1, Terminal-Bench v4 54.5; avg cost $10.79/task; avg wall time 42 min/task
Artificial Analysis Intelligence Index39.4LowClaude Opus 5 (Adaptive Reasoning, Low Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-5-low; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $1.10/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index44.8MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-5-medium; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $2.19/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index48.1HighClaude Opus 5 (Adaptive Reasoning, High Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-5-high; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $3.61/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index49.7Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-5-xhigh; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $4.88/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index50.8MaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-5; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $5.86/task (AA marks this variant deprecated)
Artificial Analysis output speed47 tok/sLowClaude Opus 5 (Adaptive Reasoning, Low Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 3.0s; list price $5/25 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed50 tok/sMediumClaude Opus 5 (Adaptive Reasoning, Medium Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 10.9s; list price $5/25 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed53 tok/sHighClaude Opus 5 (Adaptive Reasoning, High Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 22.9s; list price $5/25 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed51 tok/sExtra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 48.4s; list price $5/25 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed51 tok/sMaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 68.3s; list price $5/25 per 1M in/out (AA marks this variant deprecated)
ArXivMath (no tools)90.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—ArXivMath June 2026, no tools
ArXivMath (with tools)91.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—June 2026, with tools
AutomationBench20.4%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps); Zapier public leaderboard; vendor-reported cost/task $0.75
AutomationBench24.0%Mediumadaptive thinking, effort=mediumMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—$0.89/task
AutomationBench23.9%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps); Zapier public leaderboard; vendor-reported cost/task $0.89
AutomationBench20.5%Higheffort=highIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $2.27; competitor number as plotted by OpenAI
AutomationBench20.6%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps); Zapier public leaderboard; vendor-reported cost/task $1.03
AutomationBench25.3%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps); Zapier public leaderboard; vendor-reported cost/task $1.15
AutomationBench26.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps); Zapier public leaderboard; vendor-reported cost/task $1.27
BrowseComp90.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
Chartography (with tools)83.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
CursorBench40.7%LowOpus 5 LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 26; $4.87/task; 31,995 tokens/task; 57 steps/task
CursorBench40.7%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $4.87
CursorBench43.3%MediumOpus 5 MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 19; $6.94/task; 45,272 tokens/task; 72 steps/task
CursorBench43.3%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $6.94
CursorBench44.7%HighOpus 5 HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 16; $9.00/task; 61,405 tokens/task; 86 steps/task
CursorBench44.7%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $9.0
CursorBench46.1%Extra highOpus 5 Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 14; $11.43/task; 80,094 tokens/task; 103 steps/task
CursorBench46.1%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $11.43
CursorBench46.6%MaxOpus 5 MaxIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 12; $11.95/task; 85,384 tokens/task; 106 steps/task
CursorBench46.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $11.95
CursorBench 3.2.070.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Cursor agentCursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0
DeepSWE58.1%Loweffort=lowIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $1.66; competitor number as plotted by OpenAI
DeepSWE58.1%Lowclaude-opus-5_lowIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.0%; ±2.3; 4 runs; $1.66/task
DeepSWE68.9%Mediumclaude-opus-5_mediumIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 89.4%; ±1.2; 4 runs; $3.29/task
DeepSWE68.9%Mediumeffort=mediumIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $3.29; competitor number as plotted by OpenAI
DeepSWE72.8%Higheffort=highIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $6.08; competitor number as plotted by OpenAI
DeepSWE72.8%Highclaude-opus-5_highIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 87.6%; ±1.9; 4 runs; $6.08/task
DeepSWE73.2%Extra highclaude-opus-5_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.8%; ±3.1; 4 runs; $9.07/task
DeepSWE73.2%Extra higheffort=xhighIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $9.07; competitor number as plotted by OpenAI
DeepSWE73.7%Maxeffort=maxIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $11.84; competitor number as plotted by OpenAI
DeepSWE73.6%Maxclaude-opus-5_maxIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 88.5%; ±3.9; 4 runs; $11.84/task
DeepSWE62.5%MaxOpus 5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE68.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—DeepSWE v1.1, avg of 5 trials; read from launch-post benchmark table image (values printed in table)
Design Arena (all categories)1329Defaultclaude-opus-5Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±3.0 SE; 15954 battles; win rate 57.9%
Design Arena (fullstack)1304Defaultclaude-opus-5Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±10.7 SE; 1195 battles; win rate 57.3%
Epoch Capabilities Index162.9DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 160.2-166.7; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)2133Defaultclaude-opus-5Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 3; rubric score 17.07/20; slop 6.59; avg length 6003 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
Frontier-Bench v0.125.0%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026mini-SWE-agent (GKE)Frontier-Bench v0.1 internal run, 74 tasks, mean reward over 5 attempts; xhigh best, max "within noise, landing at 43%"; high uses 19% fewer output tokens than xhigh, low 64% fewer
Frontier-Bench v0.139.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026mini-SWE-agent (GKE)Frontier-Bench v0.1 internal run, 74 tasks, mean reward over 5 attempts; xhigh best, max "within noise, landing at 43%"; high uses 19% fewer output tokens than xhigh, low 64% fewer
Frontier-Bench v0.144.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026mini-SWE-agent (GKE)Frontier-Bench v0.1 internal run, 74 tasks, mean reward over 5 attempts; xhigh best, max "within noise, landing at 43%"; high uses 19% fewer output tokens than xhigh, low 64% fewer
Frontier-Bench v0.143.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026HarborFrontier-Bench v0.1 (agentic terminal coding); table value from Harbor evaluation; read from launch-post benchmark table image (values printed in table)
Frontier-Bench v0.143.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026mini-SWE-agent (GKE)Frontier-Bench v0.1 internal run, 74 tasks, mean reward over 5 attempts; xhigh best, max "within noise, landing at 43%"; high uses 19% fewer output tokens than xhigh, low 64% fewer
FrontierCode41.9%Loweffort=lowIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $2.68; competitor number as plotted by OpenAI
FrontierCode42.0%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Opus 5 peaks at medium; vendor-reported cost/task $2.64
FrontierCode53.4%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Opus 5 peaks at medium; vendor-reported cost/task $4.61
FrontierCode48.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Opus 5 peaks at medium; vendor-reported cost/task $7.62
FrontierCode43.6%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Opus 5 peaks at medium; vendor-reported cost/task $8.99
FrontierCode53.4%Maxclaude-opus-5_maxIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026Claude Coderead from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5
FrontierCode48.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Opus 5 peaks at medium; vendor-reported cost/task $12.28
FrontierCode v1.1 (Extended)63.6%Mediumadaptive thinking, effort=medium (best)Maker's own figureVendor-reportedAnthropic ↗24 Jul 2026Claude CodeFrontierCode v1.1 Extended, best score at medium
FrontierMath (Tiers 1-3)85.6%Maxclaude-opus-5_maxIndependent testIndependentEpoch AI ↗24 Jul 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.1pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 473.2%Maxclaude-opus-5_maxIndependent testIndependentEpoch AI ↗24 Jul 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 7.0pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierSWE v252.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026Proximal harnessreported as 0.52
GDPval-AA v21827Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—xhigh uses 25% fewer output tokens than max (1861)
GDPval-AA v21824Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—read from launch-post benchmark table image (values printed in table)
GDPval-AA v2.11294Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $0.58
GDPval-AA v2.11476Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $1.36
GDPval-AA v2.11581Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $3.03
GDPval-AA v2.11676Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $4.96
GDPval-AA v2.11708MaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1682.9-1732.95
GDPval-AA v2.11708Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $6.76
GPQA Diamond88.9%LowClaude Opus 5 (Adaptive Reasoning, Low Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-low) (AA marks this variant deprecated)
GPQA Diamond91.9%MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-medium) (AA marks this variant deprecated)
GPQA Diamond93.7%HighClaude Opus 5 (Adaptive Reasoning, High Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-high) (AA marks this variant deprecated)
GPQA Diamond93.7%Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-xhigh) (AA marks this variant deprecated)
GPQA Diamond93.9%Maxclaude-opus-5_maxIndependent testIndependentEpoch AI ↗24 Jul 2026—Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond93.2%MaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5) (AA marks this variant deprecated)
GPQA Diamond93.4%DefaultIndependent testIndependentVals.ai ↗1 Sep 2026—±1.244 stderr; $0.095/test
GPQA Diamond92.9%Defaultclaude-opus-5Independent testIndependentEpoch AI ↗6 Aug 2026—Epoch-run GPQA Diamond; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
Harvey LAB-AA93.5%MaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
HealthBench Professional59.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
HLE Diamond38.6%Defaultclaude-opus-5Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 4 (Scale rank accounts for CI); ±3; entry added 2026-09-09
Humanity's Last Exam43.4%LowClaude Opus 5 (Adaptive Reasoning, Low Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-low) (AA marks this variant deprecated)
Humanity's Last Exam51.3%MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-medium) (AA marks this variant deprecated)
Humanity's Last Exam52.8%HighClaude Opus 5 (Adaptive Reasoning, High Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-high) (AA marks this variant deprecated)
Humanity's Last Exam54.4%Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-xhigh) (AA marks this variant deprecated)
Humanity's Last Exam54.9%MaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5) (AA marks this variant deprecated)
Humanity's Last Exam56.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—no tools
Humanity's Last Exam (with tools)63.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—read from launch-post benchmark table image (values printed in table)
IOI (Vals)84.3%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±9.964 stderr; $16.477/test
LegalBench (Vals)87.0%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-opus-5; rank 8/149; ±0.417 stderr; $0.011976/test
LiveCodeBench89.0%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±0.913 stderr; $0.134/test
LLM Creative Story-Writing Benchmark (Lech Mazur)3.5HighClaude Opus 5 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 5/56; Thurstone comparison score (centered at 0); est. win chance 91%; 95% bootstrap 3.344 to 3.557
LLM Creative Story-Writing Benchmark (Lech Mazur)3.8Extra highClaude Opus 5 (xhigh)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 2/56; Thurstone comparison score (centered at 0); est. win chance 93%; 95% bootstrap 3.726 to 3.836
LMArena Code Arena (WebDev)1660Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 11 (CI rank 8-13); 95% CI 1654-1666; 21088 votes
LMArena Code Arena (WebDev)1694Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 6 (CI rank 5-8); 95% CI 1687-1700; 16747 votes
LMArena Text - Coding category1534Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 15 (CI rank 7-31); 95% CI 1527-1540; 15273 votes
LMArena Text - Coding category1530Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 21 (CI rank 7-44); 95% CI 1522-1538; 7234 votes
LMArena Text - Creative Writing1472Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 13 (rank range 7-33); 95% CI 1465.4-1479.3; 12964 votes; style-controlled
LMArena Text - Creative Writing1470Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 15 (rank range 7-35); 95% CI 1460.9-1478.2; 6376 votes; style-controlled
LMArena Text - Expert1542Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 3 (rank range 1-17); 95% CI 1533.8-1550.9; 6572 votes; style-controlled
LMArena Text - Expert1530Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 12 (rank range 1-34); 95% CI 1518.1-1541.3; 3073 votes; style-controlled
LMArena Text - Hard Prompts1516Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 11 (rank range 7-21); 95% CI 1511.2-1520.5; 38320 votes; style-controlled
LMArena Text - Hard Prompts1515Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 13 (rank range 7-24); 95% CI 1508.7-1520.4; 18450 votes; style-controlled
LMArena Text - Instruction Following1497Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 8 (rank range 4-16); 95% CI 1491.2-1502.8; 20700 votes; style-controlled
LMArena Text - Instruction Following1493Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 10 (rank range 4-21); 95% CI 1485.6-1499.8; 9911 votes; style-controlled
LMArena Text - Longer Query1505Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 10 (rank range 6-22); 95% CI 1499.3-1510.2; 27765 votes; style-controlled
LMArena Text - Longer Query1500Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 16 (rank range 7-28); 95% CI 1493.0-1506.2; 13516 votes; style-controlled
LMArena Text - Non-English1482Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 14 (rank range 5-23); 95% CI 1477.2-1486.7; 34543 votes; style-controlled
LMArena Text - Non-English1481Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 15 (rank range 5-25); 95% CI 1475.0-1486.8; 16873 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1486Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 17 (rank range 6-37); 95% CI 1479.6-1493.3; 11158 votes; style-controlled
LMArena Text - Occupational: Legal & Government1503Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 9 (rank range 1-35); 95% CI 1493.7-1512.6; 4946 votes; style-controlled
LMArena Text - Occupational: Legal & Government1503Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 10 (rank range 1-43); 95% CI 1489.6-1515.3; 2454 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1482Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 11 (rank range 7-26); 95% CI 1476.1-1488.6; 16518 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1479Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 15 (rank range 7-33); 95% CI 1471.4-1486.9; 8098 votes; style-controlled
LMArena Text (overall)1491Highclaude-opus-5-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 13 (CI rank 5-22); 95% CI 1487-1495; 58380 votes
LMArena Text (overall)1489Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 14 (CI rank 7-27); 95% CI 1484-1494; 28351 votes
MCP Atlas85.8%Extra highclaude-opus-5 (xhigh)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±2.1; entry added 2026-08-03
OfficeQA Pro66.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
OSWorld 2.022.3%Lowclaude-opus-5_lowIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.5489; tool setting batch tool; step budget 500
OSWorld 2.025.3%Mediumclaude-opus-5_mediumIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.6049; tool setting batch tool; step budget 500
OSWorld 2.029.0%Highclaude-opus-5_highIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.6379; tool setting batch tool; step budget 500
OSWorld 2.030.2%Extra highclaude-opus-5_xhighIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.6767000000000001; tool setting batch tool; step budget 500
OSWorld 2.031.4%Maxclaude-opus-5_maxIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.6831; tool setting batch tool; step budget 500
OSWorld 2.075.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—partial score; August 2026 task release; safeguard interventions scored 0
OSWorld 2.0 (strict pass rate)39.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗1 Sep 2026—strict pass rate; August 2026 task release
OSWorld 2.0 offline set55.2%Loweffort=lowIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $9.88; competitor number as plotted by OpenAI
OSWorld 2.0 offline set60.3%Mediumeffort=mediumIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $12.67; competitor number as plotted by OpenAI
OSWorld 2.0 offline set65.9%Higheffort=highIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $15.95; competitor number as plotted by OpenAI
OSWorld 2.0 offline set70.1%Extra higheffort=xhighIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $23.91; competitor number as plotted by OpenAI
OSWorld 2.0 offline set70.2%Maxeffort=maxIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗22 Sep 2026—OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $24.11; competitor number as plotted by OpenAI
OSWorld 2.1 (partial score)74.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—partial score; re-evaluated by Anthropic under Opus 5.5 conditions
OSWorld 2.1 (strict pass rate)37.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—
ProgramBench (avg test pass rate)74.7%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 4.5%, almost (>=95% tests) 37.0%; avg cost $50.53/task
ProgramBench (fully resolved)4.5%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 37.0%; avg cost $50.53/task
ProgramBench (fully resolved)3.0%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.209 stderr; $60.289/test; strict fully-resolved rate
ProgramBench (fully resolved)85.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026mini-swe-agent
SciCode49.2%LowClaude Opus 5 (Adaptive Reasoning, Low Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-low) (AA marks this variant deprecated)
SciCode51.5%MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-medium) (AA marks this variant deprecated)
SciCode55.4%HighClaude Opus 5 (Adaptive Reasoning, High Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-high) (AA marks this variant deprecated)
SciCode55.7%Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-xhigh) (AA marks this variant deprecated)
SciCode56.4%MaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5) (AA marks this variant deprecated)
SimpleBench80.6%DefaultClaude Opus 5Independent testIndependentSimpleBench ↗24 Jul 2026—AVG@5, temp 0.7; rank 8th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified59.9%Maxclaude-opus-5_maxIndependent testIndependentEpoch AI ↗10 Aug 2026—Epoch-run (no tools); ±1.55 stderr
SkillsBench60.4%Maxeffort=maxIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.578 stderr; $2.509/test
SWE Atlas - Codebase QnA63.2%Extra highOpus 5 (Claude Code) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±5.01; entry added 2026-08-04
SWE Atlas - Codebase QnA62.1%MaxOpus 5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Test Writing62.2%Extra highOpus 5 (Claude Code) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±5.58; entry added 2026-08-04
SWE-bench Multilingual89.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—
SWE-bench Multimodal59.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—
SWE-Bench Pro (public, v1)79.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—avg of 5 trials
SWE-Bench Pro V2 (full)99.4%Extra highOpus 5 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±0.4; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-Bench Pro V2 (hard)98.0%Extra highOpus 5 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-bench Verified96.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—avg of 5 trials
SWE-bench Verified97.0%DefaultIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±0.764 stderr; $1.291/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
SWE-rebench63.4%HighOpus 5 [high]Independent testIndependentSWE-rebench (Nebius) ↗1 Oct 2026SWE-rebench standard scaffoldtime window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.35; pass@5 74.8%; $3.47/problem
Terminal-Bench 2.176.4%LowClaude Opus 5 (Adaptive Reasoning, Low Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-low) (AA marks this variant deprecated)
Terminal-Bench 2.186.1%MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-medium) (AA marks this variant deprecated)
Terminal-Bench 2.187.6%HighClaude Opus 5 (Adaptive Reasoning, High Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-high) (AA marks this variant deprecated)
Terminal-Bench 2.184.6%Higheffort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.991 stderr; $0.887/test
Terminal-Bench 2.188.0%Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-xhigh) (AA marks this variant deprecated)
Terminal-Bench 2.189.1%MaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5) (AA marks this variant deprecated)
Terminal-Bench 3.041.0%Highshipped default effort (high)Maker's own figureVendor-reportedAnthropic ↗1 Oct 2026Claude Managed Agents74 tasks, two runs per model; $28 per solved task
Terminal-Bench 3.042.7%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026mini-SWE-agentTB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 7.3B, run cost $5.8k
Terminal-Bench 4.034.9%LowIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.94 95% CI; 330 trials; total run cost $2394; model release 2026-07-24
Terminal-Bench 4.026.3%LowClaude Opus 5 (Adaptive Reasoning, Low Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-low) (AA marks this variant deprecated)
Terminal-Bench 4.028.5%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; public leaderboard reports Opus 5 at 51.8%; vendor-reported cost/task $4.25
Terminal-Bench 4.044.9%MediumIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.83 95% CI; 330 trials; total run cost $3192; model release 2026-07-24
Terminal-Bench 4.034.3%MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-medium) (AA marks this variant deprecated)
Terminal-Bench 4.041.2%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; public leaderboard reports Opus 5 at 51.8%; vendor-reported cost/task $7.0
Terminal-Bench 4.050.3%HighIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.73 95% CI; 330 trials; total run cost $4662; model release 2026-07-24
Terminal-Bench 4.046.0%HighClaude Opus 5 (Adaptive Reasoning, High Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-high) (AA marks this variant deprecated)
Terminal-Bench 4.047.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; public leaderboard reports Opus 5 at 51.8%; vendor-reported cost/task $10.64
Terminal-Bench 4.053.9%Extra highIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.17 95% CI; 330 trials; total run cost $6086; model release 2026-07-24
Terminal-Bench 4.046.5%Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-xhigh) (AA marks this variant deprecated)
Terminal-Bench 4.050.6%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; public leaderboard reports Opus 5 at 51.8%; vendor-reported cost/task $13.48
Terminal-Bench 4.054.5%MaxOpus 5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.053.5%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.336 stderr; $18.597/test
Terminal-Bench 4.051.8%MaxIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.39 95% CI; 330 trials; total run cost $5969; model release 2026-07-24
Terminal-Bench 4.049.0%MaxClaude Opus 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5) (AA marks this variant deprecated)
Terminal-Bench 4.052.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; public leaderboard reports Opus 5 at 51.8%; vendor-reported cost/task $15.83
Terminal-Bench-Science 0.129.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); public leaderboard: 30.0%
Toolathlon80.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—internal harness, Pass@1
Vals Code Migration57.5%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.371 stderr; $60.510/test
Vals CorpFin v273.2%Defaultvals id anthropic/claude-opus-5Independent testIndependentVals.ai ↗12 Aug 2026—vals id anthropic/claude-opus-5; rank 1/134; ±0.871 stderr; $0.860452/test
Vals Index63.7%Maxeffort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±0.957 stderr; $19.315/test
Vals Legal Research Bench55.3%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-opus-5; rank 2/72; ±3.456 stderr; $6.762498/test
Vals Public Benefits Bench v1.176.9%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-opus-5; rank 1/45; ±1.096 stderr; $3.949099/test
Vals SRE Bench12.2%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±2.027 stderr; $23.633/test
Vals TaxEval v275.1%Defaultvals id anthropic/claude-opus-5Independent testIndependentVals.ai ↗1 Sep 2026—vals id anthropic/claude-opus-5; rank 23/145; ±0.832 stderr; $0.114274/test
Vals Vibe Code Bench88.4%DefaultIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±3.001 stderr; $33.876/test
Vending-Bench 2$11,182DefaultIndependent testIndependentAndon Labs ↗1 Oct 2026—final money balance after simulated year, arithmetic mean across runs; ±$2,094; rank 4; only top 10 rendered server-side (57 more behind 'Show more')
WANDR50.5%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $10.65
WANDR58.1%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $23.97
WANDR64.6%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $43.49
WANDR67.0%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $53.84
WANDR67.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $61.62