Models · Anthropic · Out sinceReleased 24 Jul 2026
Claude Opus 5#2 for research and analysis.#2 for research and analysis, best at max effort.
Claude Opus 5 is made by Anthropic. Among the models we track it ranks #2 for research and analysis, #6 for writing, #6 for coding. It's expensive to use.
API id claude-opus-5. Default effort high; thinking cannot be disabled at xhigh/max. Previous-generation Opus.
- 71.1 / 100 · #6
- 84.7 / 100 · #2
- 80.5 / 100 · #6
- Expensive$5 / $25
- $10
- 1M
- 128K
- 222 (153 independent153 indep.)
- 24 Jul 2026
- Anthropic
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for research and analysis (Maximum thinking).
Highlighted: dominant setting in its research and analysis composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-opus-5.
Route it as anthropic/claude-opus-5 at $5 in / $25 out per 1M tokens, 1M context. Listed since 24 Jul 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolsverbosity
Thinking level
Reasoning effort
How long should Claude Opus 5Where Claude Opus 5think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Artificial Analysis Coding Agent Index
Best at Extra high, worse abovePeaks at Extra high
Terminal-Bench 4.0
Best at MaxBest at Max
CursorBench
Best at MaxBest at Max
DeepSWE
Best at MaxBest at Max
FrontierCode
Best at MediumBest at Medium
Frontier-Bench v0.1
Best at Extra high, worse abovePeaks at Extra high
LMArena Code Arena (WebDev)
Best at MaxBest at Max
SWE Atlas - Codebase QnA
Best at Extra high, worse abovePeaks at Extra high
ProgramBench (fully resolved)
Best at Extra high, worse abovePeaks at Extra high
Terminal-Bench 3.0
Best at MaxBest at Max
Terminal-Bench 2.1
Best at MaxBest at Max
LMArena Text - Coding category
Best at High, worse abovePeaks at High
SciCode
Best at MaxBest at Max
Agents' Last Exam
Best at High, worse abovePeaks at High
Artificial Analysis Intelligence Index
Best at MaxBest at Max
AutomationBench
Best at MaxBest at Max
OSWorld 2.0
Best at MaxBest at Max
OSWorld 2.0 offline set
Best at MaxBest at Max
AA-Briefcase
Best at MaxBest at Max
ARC-AGI-2
Best at MaxBest at Max
GDPval-AA v2
Best at Extra high, worse abovePeaks at Extra high
GDPval-AA v2.1
Best at MaxBest at Max
GPQA Diamond
Best at High, worse abovePeaks at High
Humanity's Last Exam
Best at MaxBest at Max
LLM Creative Story-Writing Benchmark (Lech Mazur)
Best at Extra highBest at Extra high
LMArena Text - Creative Writing
Best at High, worse abovePeaks at High
LMArena Text - Expert
Best at High, worse abovePeaks at High
LMArena Text - Hard Prompts
Best at High, worse abovePeaks at High
LMArena Text - Instruction Following
Best at High, worse abovePeaks at High
LMArena Text - Longer Query
Best at High, worse abovePeaks at High
LMArena Text - Non-English
Best at High, worse abovePeaks at High
LMArena Text - Occupational: Legal & Government
Best at High, worse abovePeaks at High
LMArena Text - Occupational: Writing, Literature & Language
Best at High, worse abovePeaks at High
LMArena Text (overall)
Best at High, worse abovePeaks at High
WANDR
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA Analyst Agent | 53.8% | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run; scores move in 1.25-pt steps (small task set) |
| AA-Briefcase | 1606 | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | AA-Briefcase (pre-v1.1), run by Artificial Analysis; xhigh/high use 15%/33% fewer output tokens than max |
| AA-Briefcase | 1693 | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | AA-Briefcase (pre-v1.1), run by Artificial Analysis; xhigh/high use 15%/33% fewer output tokens than max |
| AA-Briefcase | 1720 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | AA-Briefcase (pre-v1.1), run by Artificial Analysis; xhigh/high use 15%/33% fewer output tokens than max |
| AA-Briefcase v1.1 | 1673 | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Briefcase v1.1 | 1673 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| Agents' Last Exam | 51.9% | Loweffort=low | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $4.0619; competitor number as plotted by OpenAI |
| Agents' Last Exam | 53.0% | Mediumeffort=medium | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $5.2907; competitor number as plotted by OpenAI |
| Agents' Last Exam | 55.9% | Higheffort=high | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $7.2917; competitor number as plotted by OpenAI |
| Agents' Last Exam | 55.5% | Extra higheffort=xhigh | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $10.0226; competitor number as plotted by OpenAI |
| Agents' Last Exam | 52.7% | Maxeffort=max | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | Agents' Last Exam V1 (long-horizon professional workflows, 55 sub-industries); vendor-estimated API cost/task $9.7583; competitor number as plotted by OpenAI |
| ARC-AGI-1 | 97.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | |
| ARC-AGI-2 | 88.3% | HighClaude Opus 5 (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $1.45/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 90.4% | MaxClaude Opus 5 (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $2.06/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 90.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | semi-private set |
| ARC-AGI-3 | 30.2% | HighClaude Opus 5 (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $20,657; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 30.2% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | ARC Prize Foundation verified, semi-private set (RHAE 30.16%); max-effort result not available at release |
| Artificial Analysis Coding Agent Index | 59.2 | No reasoningeffort=none | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.53; competitor number as plotted by OpenAI |
| Artificial Analysis Coding Agent Index | 59.4 | Loweffort=low | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $2.3; competitor number as plotted by OpenAI |
| Artificial Analysis Coding Agent Index | 64.1 | Mediumeffort=medium | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.17; competitor number as plotted by OpenAI |
| Artificial Analysis Coding Agent Index | 65.6 | Higheffort=high | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $3.92; competitor number as plotted by OpenAI |
| Artificial Analysis Coding Agent Index | 68.1 | Extra higheffort=xhigh | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $8.17; competitor number as plotted by OpenAI |
| Artificial Analysis Coding Agent Index | 67 | Maxeffort=max | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 3 Sep 2026 | — | Artificial Analysis Coding Agent Index v1.4 (index score 0-100); vendor-estimated API cost/task $8.94; competitor number as plotted by OpenAI |
| Artificial Analysis Coding Agent Index | 59.7 | MaxOpus 5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | agent Claude Code; components: DeepSWE v1.1 62.5, SWE-Atlas-QnA 62.1, Terminal-Bench v4 54.5; avg cost $10.79/task; avg wall time 42 min/task |
| Artificial Analysis Intelligence Index | 39.4 | LowClaude Opus 5 (Adaptive Reasoning, Low Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-5-low; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $1.10/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 44.8 | MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-5-medium; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $2.19/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 48.1 | HighClaude Opus 5 (Adaptive Reasoning, High Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-5-high; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $3.61/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 49.7 | Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-5-xhigh; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $4.88/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 50.8 | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-5; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $5.86/task (AA marks this variant deprecated) |
| Artificial Analysis output speed | 47 tok/s | LowClaude Opus 5 (Adaptive Reasoning, Low Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 3.0s; list price $5/25 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 50 tok/s | MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 10.9s; list price $5/25 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 53 tok/s | HighClaude Opus 5 (Adaptive Reasoning, High Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 22.9s; list price $5/25 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 51 tok/s | Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 48.4s; list price $5/25 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 51 tok/s | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 68.3s; list price $5/25 per 1M in/out (AA marks this variant deprecated) |
| ArXivMath (no tools) | 90.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | ArXivMath June 2026, no tools |
| ArXivMath (with tools) | 91.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | June 2026, with tools |
| AutomationBench | 20.4% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps); Zapier public leaderboard; vendor-reported cost/task $0.75 |
| AutomationBench | 24.0% | Mediumadaptive thinking, effort=medium | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | $0.89/task |
| AutomationBench | 23.9% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps); Zapier public leaderboard; vendor-reported cost/task $0.89 |
| AutomationBench | 20.5% | Higheffort=high | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $2.27; competitor number as plotted by OpenAI |
| AutomationBench | 20.6% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps); Zapier public leaderboard; vendor-reported cost/task $1.03 |
| AutomationBench | 25.3% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps); Zapier public leaderboard; vendor-reported cost/task $1.15 |
| AutomationBench | 26.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps); Zapier public leaderboard; vendor-reported cost/task $1.27 |
| BrowseComp | 90.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| Chartography (with tools) | 83.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| CursorBench | 40.7% | LowOpus 5 Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 26; $4.87/task; 31,995 tokens/task; 57 steps/task |
| CursorBench | 40.7% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $4.87 |
| CursorBench | 43.3% | MediumOpus 5 Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 19; $6.94/task; 45,272 tokens/task; 72 steps/task |
| CursorBench | 43.3% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $6.94 |
| CursorBench | 44.7% | HighOpus 5 High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 16; $9.00/task; 61,405 tokens/task; 86 steps/task |
| CursorBench | 44.7% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $9.0 |
| CursorBench | 46.1% | Extra highOpus 5 Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 14; $11.43/task; 80,094 tokens/task; 103 steps/task |
| CursorBench | 46.1% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $11.43 |
| CursorBench | 46.6% | MaxOpus 5 Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 12; $11.95/task; 85,384 tokens/task; 106 steps/task |
| CursorBench | 46.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $11.95 |
| CursorBench 3.2.0 | 70.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Cursor agent | CursorBench 3.2.0 (Cursor production harness); not comparable to CursorBench 4.0 |
| DeepSWE | 58.1% | Loweffort=low | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $1.66; competitor number as plotted by OpenAI |
| DeepSWE | 58.1% | Lowclaude-opus-5_low | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.0%; ±2.3; 4 runs; $1.66/task |
| DeepSWE | 68.9% | Mediumclaude-opus-5_medium | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 89.4%; ±1.2; 4 runs; $3.29/task |
| DeepSWE | 68.9% | Mediumeffort=medium | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $3.29; competitor number as plotted by OpenAI |
| DeepSWE | 72.8% | Higheffort=high | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $6.08; competitor number as plotted by OpenAI |
| DeepSWE | 72.8% | Highclaude-opus-5_high | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 87.6%; ±1.9; 4 runs; $6.08/task |
| DeepSWE | 73.2% | Extra highclaude-opus-5_xhigh | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.8%; ±3.1; 4 runs; $9.07/task |
| DeepSWE | 73.2% | Extra higheffort=xhigh | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $9.07; competitor number as plotted by OpenAI |
| DeepSWE | 73.7% | Maxeffort=max | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | DeepSWE v1.1 (113 long-horizon SWE tasks); GPT models run in Codex; vendor-estimated API cost/task $11.84; competitor number as plotted by OpenAI |
| DeepSWE | 73.6% | Maxclaude-opus-5_max | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 88.5%; ±3.9; 4 runs; $11.84/task |
| DeepSWE | 62.5% | MaxOpus 5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 68.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | DeepSWE v1.1, avg of 5 trials; read from launch-post benchmark table image (values printed in table) |
| Design Arena (all categories) | 1329 | Defaultclaude-opus-5 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±3.0 SE; 15954 battles; win rate 57.9% |
| Design Arena (fullstack) | 1304 | Defaultclaude-opus-5 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±10.7 SE; 1195 battles; win rate 57.3% |
| Epoch Capabilities Index | 162.9 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 160.2-166.7; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 2133 | Defaultclaude-opus-5 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 3; rubric score 17.07/20; slop 6.59; avg length 6003 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| Frontier-Bench v0.1 | 25.0% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | mini-SWE-agent (GKE) | Frontier-Bench v0.1 internal run, 74 tasks, mean reward over 5 attempts; xhigh best, max "within noise, landing at 43%"; high uses 19% fewer output tokens than xhigh, low 64% fewer |
| Frontier-Bench v0.1 | 39.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | mini-SWE-agent (GKE) | Frontier-Bench v0.1 internal run, 74 tasks, mean reward over 5 attempts; xhigh best, max "within noise, landing at 43%"; high uses 19% fewer output tokens than xhigh, low 64% fewer |
| Frontier-Bench v0.1 | 44.4% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | mini-SWE-agent (GKE) | Frontier-Bench v0.1 internal run, 74 tasks, mean reward over 5 attempts; xhigh best, max "within noise, landing at 43%"; high uses 19% fewer output tokens than xhigh, low 64% fewer |
| Frontier-Bench v0.1 | 43.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | Harbor | Frontier-Bench v0.1 (agentic terminal coding); table value from Harbor evaluation; read from launch-post benchmark table image (values printed in table) |
| Frontier-Bench v0.1 | 43.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | mini-SWE-agent (GKE) | Frontier-Bench v0.1 internal run, 74 tasks, mean reward over 5 attempts; xhigh best, max "within noise, landing at 43%"; high uses 19% fewer output tokens than xhigh, low 64% fewer |
| FrontierCode | 41.9% | Loweffort=low | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | FrontierCode 1.1 Main (Cognition); GPT models run in Codex CLI; vendor-estimated API cost/task $2.68; competitor number as plotted by OpenAI |
| FrontierCode | 42.0% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Opus 5 peaks at medium; vendor-reported cost/task $2.64 |
| FrontierCode | 53.4% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Opus 5 peaks at medium; vendor-reported cost/task $4.61 |
| FrontierCode | 48.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Opus 5 peaks at medium; vendor-reported cost/task $7.62 |
| FrontierCode | 43.6% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Opus 5 peaks at medium; vendor-reported cost/task $8.99 |
| FrontierCode | 53.4% | Maxclaude-opus-5_max | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | Claude Code | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5 |
| FrontierCode | 48.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Opus 5 peaks at medium; vendor-reported cost/task $12.28 |
| FrontierCode v1.1 (Extended) | 63.6% | Mediumadaptive thinking, effort=medium (best) | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | Claude Code | FrontierCode v1.1 Extended, best score at medium |
| FrontierMath (Tiers 1-3) | 85.6% | Maxclaude-opus-5_max | Independent testIndependentEpoch AI ↗ | 24 Jul 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.1pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 4 | 73.2% | Maxclaude-opus-5_max | Independent testIndependentEpoch AI ↗ | 24 Jul 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 7.0pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierSWE v2 | 52.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | Proximal harness | reported as 0.52 |
| GDPval-AA v2 | 1827 | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | xhigh uses 25% fewer output tokens than max (1861) |
| GDPval-AA v2 | 1824 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | read from launch-post benchmark table image (values printed in table) |
| GDPval-AA v2.1 | 1294 | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $0.58 |
| GDPval-AA v2.1 | 1476 | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $1.36 |
| GDPval-AA v2.1 | 1581 | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $3.03 |
| GDPval-AA v2.1 | 1676 | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $4.96 |
| GDPval-AA v2.1 | 1708 | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1682.9-1732.95 |
| GDPval-AA v2.1 | 1708 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $6.76 |
| GPQA Diamond | 88.9% | LowClaude Opus 5 (Adaptive Reasoning, Low Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-low) (AA marks this variant deprecated) |
| GPQA Diamond | 91.9% | MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-medium) (AA marks this variant deprecated) |
| GPQA Diamond | 93.7% | HighClaude Opus 5 (Adaptive Reasoning, High Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-high) (AA marks this variant deprecated) |
| GPQA Diamond | 93.7% | Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-xhigh) (AA marks this variant deprecated) |
| GPQA Diamond | 93.9% | Maxclaude-opus-5_max | Independent testIndependentEpoch AI ↗ | 24 Jul 2026 | — | Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 93.2% | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5) (AA marks this variant deprecated) |
| GPQA Diamond | 93.4% | Default | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.244 stderr; $0.095/test |
| GPQA Diamond | 92.9% | Defaultclaude-opus-5 | Independent testIndependentEpoch AI ↗ | 6 Aug 2026 | — | Epoch-run GPQA Diamond; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| Harvey LAB-AA | 93.5% | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| HealthBench Professional | 59.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | read from launch-post benchmark table image (values printed in table) |
| HLE Diamond | 38.6% | Defaultclaude-opus-5 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 4 (Scale rank accounts for CI); ±3; entry added 2026-09-09 |
| Humanity's Last Exam | 43.4% | LowClaude Opus 5 (Adaptive Reasoning, Low Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-low) (AA marks this variant deprecated) |
| Humanity's Last Exam | 51.3% | MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-medium) (AA marks this variant deprecated) |
| Humanity's Last Exam | 52.8% | HighClaude Opus 5 (Adaptive Reasoning, High Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-high) (AA marks this variant deprecated) |
| Humanity's Last Exam | 54.4% | Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-xhigh) (AA marks this variant deprecated) |
| Humanity's Last Exam | 54.9% | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5) (AA marks this variant deprecated) |
| Humanity's Last Exam | 56.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | no tools |
| Humanity's Last Exam (with tools) | 63.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | read from launch-post benchmark table image (values printed in table) |
| IOI (Vals) | 84.3% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±9.964 stderr; $16.477/test |
| LegalBench (Vals) | 87.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-opus-5; rank 8/149; ±0.417 stderr; $0.011976/test |
| LiveCodeBench | 89.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.913 stderr; $0.134/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 3.5 | HighClaude Opus 5 (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 5/56; Thurstone comparison score (centered at 0); est. win chance 91%; 95% bootstrap 3.344 to 3.557 |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 3.8 | Extra highClaude Opus 5 (xhigh) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 2/56; Thurstone comparison score (centered at 0); est. win chance 93%; 95% bootstrap 3.726 to 3.836 |
| LMArena Code Arena (WebDev) | 1660 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 11 (CI rank 8-13); 95% CI 1654-1666; 21088 votes |
| LMArena Code Arena (WebDev) | 1694 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 6 (CI rank 5-8); 95% CI 1687-1700; 16747 votes |
| LMArena Text - Coding category | 1534 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 15 (CI rank 7-31); 95% CI 1527-1540; 15273 votes |
| LMArena Text - Coding category | 1530 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 21 (CI rank 7-44); 95% CI 1522-1538; 7234 votes |
| LMArena Text - Creative Writing | 1472 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 13 (rank range 7-33); 95% CI 1465.4-1479.3; 12964 votes; style-controlled |
| LMArena Text - Creative Writing | 1470 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 15 (rank range 7-35); 95% CI 1460.9-1478.2; 6376 votes; style-controlled |
| LMArena Text - Expert | 1542 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 3 (rank range 1-17); 95% CI 1533.8-1550.9; 6572 votes; style-controlled |
| LMArena Text - Expert | 1530 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 12 (rank range 1-34); 95% CI 1518.1-1541.3; 3073 votes; style-controlled |
| LMArena Text - Hard Prompts | 1516 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 11 (rank range 7-21); 95% CI 1511.2-1520.5; 38320 votes; style-controlled |
| LMArena Text - Hard Prompts | 1515 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 13 (rank range 7-24); 95% CI 1508.7-1520.4; 18450 votes; style-controlled |
| LMArena Text - Instruction Following | 1497 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 8 (rank range 4-16); 95% CI 1491.2-1502.8; 20700 votes; style-controlled |
| LMArena Text - Instruction Following | 1493 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 10 (rank range 4-21); 95% CI 1485.6-1499.8; 9911 votes; style-controlled |
| LMArena Text - Longer Query | 1505 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 10 (rank range 6-22); 95% CI 1499.3-1510.2; 27765 votes; style-controlled |
| LMArena Text - Longer Query | 1500 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 16 (rank range 7-28); 95% CI 1493.0-1506.2; 13516 votes; style-controlled |
| LMArena Text - Non-English | 1482 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 14 (rank range 5-23); 95% CI 1477.2-1486.7; 34543 votes; style-controlled |
| LMArena Text - Non-English | 1481 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 15 (rank range 5-25); 95% CI 1475.0-1486.8; 16873 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1486 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 17 (rank range 6-37); 95% CI 1479.6-1493.3; 11158 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1503 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 9 (rank range 1-35); 95% CI 1493.7-1512.6; 4946 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1503 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 10 (rank range 1-43); 95% CI 1489.6-1515.3; 2454 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1482 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 11 (rank range 7-26); 95% CI 1476.1-1488.6; 16518 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1479 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 15 (rank range 7-33); 95% CI 1471.4-1486.9; 8098 votes; style-controlled |
| LMArena Text (overall) | 1491 | Highclaude-opus-5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 13 (CI rank 5-22); 95% CI 1487-1495; 58380 votes |
| LMArena Text (overall) | 1489 | Maxclaude-opus-5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 14 (CI rank 7-27); 95% CI 1484-1494; 28351 votes |
| MCP Atlas | 85.8% | Extra highclaude-opus-5 (xhigh) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±2.1; entry added 2026-08-03 |
| OfficeQA Pro | 66.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| OSWorld 2.0 | 22.3% | Lowclaude-opus-5_low | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.5489; tool setting batch tool; step budget 500 |
| OSWorld 2.0 | 25.3% | Mediumclaude-opus-5_medium | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.6049; tool setting batch tool; step budget 500 |
| OSWorld 2.0 | 29.0% | Highclaude-opus-5_high | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.6379; tool setting batch tool; step budget 500 |
| OSWorld 2.0 | 30.2% | Extra highclaude-opus-5_xhigh | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.6767000000000001; tool setting batch tool; step budget 500 |
| OSWorld 2.0 | 31.4% | Maxclaude-opus-5_max | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.6831; tool setting batch tool; step budget 500 |
| OSWorld 2.0 | 75.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | partial score; August 2026 task release; safeguard interventions scored 0 |
| OSWorld 2.0 (strict pass rate) | 39.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 1 Sep 2026 | — | strict pass rate; August 2026 task release |
| OSWorld 2.0 offline set | 55.2% | Loweffort=low | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $9.88; competitor number as plotted by OpenAI |
| OSWorld 2.0 offline set | 60.3% | Mediumeffort=medium | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $12.67; competitor number as plotted by OpenAI |
| OSWorld 2.0 offline set | 65.9% | Higheffort=high | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $15.95; competitor number as plotted by OpenAI |
| OSWorld 2.0 offline set | 70.1% | Extra higheffort=xhigh | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $23.91; competitor number as plotted by OpenAI |
| OSWorld 2.0 offline set | 70.2% | Maxeffort=max | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 22 Sep 2026 | — | OSWorld 2.0 offline set, partial reward, v2026.08.08 release; vendor-estimated API cost/task $24.11; competitor number as plotted by OpenAI |
| OSWorld 2.1 (partial score) | 74.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | partial score; re-evaluated by Anthropic under Opus 5.5 conditions |
| OSWorld 2.1 (strict pass rate) | 37.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | |
| ProgramBench (avg test pass rate) | 74.7% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 4.5%, almost (>=95% tests) 37.0%; avg cost $50.53/task |
| ProgramBench (fully resolved) | 4.5% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 37.0%; avg cost $50.53/task |
| ProgramBench (fully resolved) | 3.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.209 stderr; $60.289/test; strict fully-resolved rate |
| ProgramBench (fully resolved) | 85.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | mini-swe-agent | |
| SciCode | 49.2% | LowClaude Opus 5 (Adaptive Reasoning, Low Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-low) (AA marks this variant deprecated) |
| SciCode | 51.5% | MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-medium) (AA marks this variant deprecated) |
| SciCode | 55.4% | HighClaude Opus 5 (Adaptive Reasoning, High Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-high) (AA marks this variant deprecated) |
| SciCode | 55.7% | Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-xhigh) (AA marks this variant deprecated) |
| SciCode | 56.4% | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5) (AA marks this variant deprecated) |
| SimpleBench | 80.6% | DefaultClaude Opus 5 | Independent testIndependentSimpleBench ↗ | 24 Jul 2026 | — | AVG@5, temp 0.7; rank 8th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 59.9% | Maxclaude-opus-5_max | Independent testIndependentEpoch AI ↗ | 10 Aug 2026 | — | Epoch-run (no tools); ±1.55 stderr |
| SkillsBench | 60.4% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.578 stderr; $2.509/test |
| SWE Atlas - Codebase QnA | 63.2% | Extra highOpus 5 (Claude Code) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 1 (Scale rank accounts for CI); ±5.01; entry added 2026-08-04 |
| SWE Atlas - Codebase QnA | 62.1% | MaxOpus 5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Test Writing | 62.2% | Extra highOpus 5 (Claude Code) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 1 (Scale rank accounts for CI); ±5.58; entry added 2026-08-04 |
| SWE-bench Multilingual | 89.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | |
| SWE-bench Multimodal | 59.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | |
| SWE-Bench Pro (public, v1) | 79.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | avg of 5 trials |
| SWE-Bench Pro V2 (full) | 99.4% | Extra highOpus 5 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 1 (Scale rank accounts for CI); ±0.4; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-Bench Pro V2 (hard) | 98.0% | Extra highOpus 5 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Claude Code | rank 1 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-bench Verified | 96.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 24 Jul 2026 | — | avg of 5 trials |
| SWE-bench Verified | 97.0% | Default | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±0.764 stderr; $1.291/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| SWE-rebench | 63.4% | HighOpus 5 [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | 1 Oct 2026 | SWE-rebench standard scaffold | time window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.35; pass@5 74.8%; $3.47/problem |
| Terminal-Bench 2.1 | 76.4% | LowClaude Opus 5 (Adaptive Reasoning, Low Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-low) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 86.1% | MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-medium) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 87.6% | HighClaude Opus 5 (Adaptive Reasoning, High Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-high) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 84.6% | Higheffort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±0.991 stderr; $0.887/test |
| Terminal-Bench 2.1 | 88.0% | Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-xhigh) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 89.1% | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5) (AA marks this variant deprecated) |
| Terminal-Bench 3.0 | 41.0% | Highshipped default effort (high) | Maker's own figureVendor-reportedAnthropic ↗ | 1 Oct 2026 | Claude Managed Agents | 74 tasks, two runs per model; $28 per solved task |
| Terminal-Bench 3.0 | 42.7% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | mini-SWE-agent | TB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 7.3B, run cost $5.8k |
| Terminal-Bench 4.0 | 34.9% | Low | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.94 95% CI; 330 trials; total run cost $2394; model release 2026-07-24 |
| Terminal-Bench 4.0 | 26.3% | LowClaude Opus 5 (Adaptive Reasoning, Low Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-low) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 28.5% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; public leaderboard reports Opus 5 at 51.8%; vendor-reported cost/task $4.25 |
| Terminal-Bench 4.0 | 44.9% | Medium | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.83 95% CI; 330 trials; total run cost $3192; model release 2026-07-24 |
| Terminal-Bench 4.0 | 34.3% | MediumClaude Opus 5 (Adaptive Reasoning, Medium Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-medium) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 41.2% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; public leaderboard reports Opus 5 at 51.8%; vendor-reported cost/task $7.0 |
| Terminal-Bench 4.0 | 50.3% | High | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.73 95% CI; 330 trials; total run cost $4662; model release 2026-07-24 |
| Terminal-Bench 4.0 | 46.0% | HighClaude Opus 5 (Adaptive Reasoning, High Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-high) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 47.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; public leaderboard reports Opus 5 at 51.8%; vendor-reported cost/task $10.64 |
| Terminal-Bench 4.0 | 53.9% | Extra high | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.17 95% CI; 330 trials; total run cost $6086; model release 2026-07-24 |
| Terminal-Bench 4.0 | 46.5% | Extra highClaude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-xhigh) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 50.6% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; public leaderboard reports Opus 5 at 51.8%; vendor-reported cost/task $13.48 |
| Terminal-Bench 4.0 | 54.5% | MaxOpus 5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 53.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.336 stderr; $18.597/test |
| Terminal-Bench 4.0 | 51.8% | Max | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | Claude Code | TB 4.0.0 (66 tasks); ±3.39 95% CI; 330 trials; total run cost $5969; model release 2026-07-24 |
| Terminal-Bench 4.0 | 49.0% | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 52.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; public leaderboard reports Opus 5 at 51.8%; vendor-reported cost/task $15.83 |
| Terminal-Bench-Science 0.1 | 29.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench-Science 0.1 (70 tasks); public leaderboard: 30.0% |
| Toolathlon | 80.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | internal harness, Pass@1 |
| Vals Code Migration | 57.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.371 stderr; $60.510/test |
| Vals CorpFin v2 | 73.2% | Defaultvals id anthropic/claude-opus-5 | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id anthropic/claude-opus-5; rank 1/134; ±0.871 stderr; $0.860452/test |
| Vals Index | 63.7% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±0.957 stderr; $19.315/test |
| Vals Legal Research Bench | 55.3% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-opus-5; rank 2/72; ±3.456 stderr; $6.762498/test |
| Vals Public Benefits Bench v1.1 | 76.9% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-opus-5; rank 1/45; ±1.096 stderr; $3.949099/test |
| Vals SRE Bench | 12.2% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±2.027 stderr; $23.633/test |
| Vals TaxEval v2 | 75.1% | Defaultvals id anthropic/claude-opus-5 | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id anthropic/claude-opus-5; rank 23/145; ±0.832 stderr; $0.114274/test |
| Vals Vibe Code Bench | 88.4% | Default | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±3.001 stderr; $33.876/test |
| Vending-Bench 2 | $11,182 | Default | Independent testIndependentAndon Labs ↗ | 1 Oct 2026 | — | final money balance after simulated year, arithmetic mean across runs; ±$2,094; rank 4; only top 10 rendered server-side (57 more behind 'Show more') |
| WANDR | 50.5% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $10.65 |
| WANDR | 58.1% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $23.97 |
| WANDR | 64.6% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $43.49 |
| WANDR | 67.0% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $53.84 |
| WANDR | 67.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $61.62 |