Models · Anthropic · Out sinceReleased 22 Sep 2026
Claude Opus 5.5#2 for writing.#2 for writing, best at high effort.
Claude Opus 5.5 is made by Anthropic. Among the models we track it ranks #2 for writing, #2 for coding, #4 for research and analysis. It's expensive to use. In blind tests where thousands of people compare two anonymous texts, people prefer Claude Opus 5.5's writing over every other model you can use today. It's also the strongest at long, well-built stories and texts.
API id claude-opus-5-5. Default effort = medium (unlike other Claude models, which default to high). Adaptive thinking always on (thinking cannot be disabled). Cache reads $0.20/M, cache writes $5/M. Fast mode (up to 2.5x speed) $8/$40. Cyber/bio safeguards with fallback to Opus 4.8 / Opus 5. Knowledge cutoff Jun 2026.
- 86.6 / 100 · #2
- 81.2 / 100 · #4
- 89.9 / 100 · #2
- Expensive$4 / $20
- $8
- 1M
- 128K
- 168 (110 independent110 indep.)
- 22 Sep 2026
- Anthropic
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (High thinking).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-opus-5.5.
Route it as anthropic/claude-opus-5.5 at $4 in / $20 out per 1M tokens, 1M context. Listed since 22 Sep 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolsverbosity
Thinking level
Reasoning effort
How long should Claude Opus 5.5Where Claude Opus 5.5think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at MaxBest at Max
CursorBench
Best at MaxBest at Max
FrontierCode
Best at Medium, worse abovePeaks at Medium
FrontierCode v1.1 (Extended)
Best at Medium, worse abovePeaks at Medium
SciCode
Best at MaxBest at Max
SWE-bench Pro (Anthropic internal subset)
Best at HighBest at High
Artificial Analysis Intelligence Index
Best at MaxBest at Max
AutomationBench
Best at MaxBest at Max
AA-Briefcase v1.1
Best at MaxBest at Max
AA-LCR
Best at Extra highBest at Extra high
AA-Omniscience Accuracy
Best at MaxBest at Max
AA-Omniscience Hallucination Rate
Best at MaxBest at Max
Lower is better on this test
AA-Omniscience Index
Best at MaxBest at Max
ARC-AGI-2
Best at High, worse abovePeaks at High
GDP.pdf
Best at High, worse abovePeaks at High
GDPval-AA v2.1
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
WANDR
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1285 | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $1.15 |
| AA-Briefcase v1.1 | 1642 | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $4.4 |
| AA-Briefcase v1.1 | 1705 | HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Briefcase v1.1 | 1705 | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $6.27 |
| AA-Briefcase v1.1 | 1780 | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Briefcase v1.1 | 1780 | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $12.27 |
| AA-Briefcase v1.1 | 1822 | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Briefcase v1.1 | 1822 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $21.05 |
| AA-LCR | 80.7% | LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 84.3% | MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 82.7% | HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 84.7% | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 84.7% | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 63.5% | LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 64.5% | MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 64.6% | HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 65.4% | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 66.2% | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 67.6% | LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 68.4% | MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 67.6% | HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 65.7% | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 58.6% | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 38.9 | LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 40.3 | MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 40.6 | HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 42.7 | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 46.4 | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| ARC-AGI-2 | 87.5% | MediumClaude Opus 5.5 (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.34/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 93.3% | HighClaude Opus 5.5 (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.41/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 92.5% | Extra highClaude Opus 5.5 (XHigh) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.67/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-2 | 91.7% | MaxClaude Opus 5.5 (Max) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $1.85/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| Artificial Analysis Coding Agent Index | 66 | MaxOpus 5.5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | agent Claude Code; components: DeepSWE v1.1 68.4, SWE-Atlas-QnA 66.4, Terminal-Bench v4 63.1; avg cost $13.04/task; avg wall time 64 min/task |
| Artificial Analysis Intelligence Index | 42.3 | LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-5-5-low; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $0.55/task |
| Artificial Analysis Intelligence Index | 51.2 | MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-5-5-medium; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $1.34/task |
| Artificial Analysis Intelligence Index | 53.6 | HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-5-5-high; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $1.82/task |
| Artificial Analysis Intelligence Index | 56 | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-5-5-xhigh; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $3.46/task |
| Artificial Analysis Intelligence Index | 57.6 | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-opus-5-5; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $5.98/task |
| Artificial Analysis output speed | 72 tok/s | LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 13.5s; list price $4/20 per 1M in/out |
| Artificial Analysis output speed | 71 tok/s | MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 24.4s; list price $4/20 per 1M in/out |
| Artificial Analysis output speed | 73 tok/s | HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 38.9s; list price $4/20 per 1M in/out |
| Artificial Analysis output speed | 77 tok/s | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 137.2s; list price $4/20 per 1M in/out |
| Artificial Analysis output speed | 90 tok/s | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 724.5s; list price $4/20 per 1M in/out |
| ArXivMath (no tools) | 91.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | ArXivMath Aug 2026 (57 problems), no tools, avg of 4 |
| ArXivMath (with tools) | 96.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | Aug 2026, code execution |
| AutomationBench | 24.2% | Loweffort=low | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.51; competitor number as plotted by OpenAI ("Opus 5.5 w/ fallbacks") |
| AutomationBench | 23.3% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps); run by Zapier without fallback models (safeguard interventions counted as failures); vendor-reported cost/task $0.5 |
| AutomationBench | 29.5% | Mediumeffort=medium | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.65; competitor number as plotted by OpenAI ("Opus 5.5 w/ fallbacks") |
| AutomationBench | 28.6% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps); run by Zapier without fallback models (safeguard interventions counted as failures); vendor-reported cost/task $0.64 |
| AutomationBench | 33.0% | Higheffort=high | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.71; competitor number as plotted by OpenAI ("Opus 5.5 w/ fallbacks") |
| AutomationBench | 32.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps); run by Zapier without fallback models (safeguard interventions counted as failures); vendor-reported cost/task $0.7 |
| AutomationBench | 35.8% | Extra higheffort=xhigh | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 29 Sep 2026 | — | AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.89; competitor number as plotted by OpenAI ("Opus 5.5 w/ fallbacks") |
| AutomationBench | 34.4% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps); run by Zapier without fallback models (safeguard interventions counted as failures); vendor-reported cost/task $0.86 |
| AutomationBench | 40.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | AutomationBench (Zapier; end-to-end business workflows across connected apps); run by Zapier without fallback models (safeguard interventions counted as failures); vendor-reported cost/task $1.37 |
| Chartography (no tools) | 64.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | no tools |
| Chartography (with tools) | 89.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | Chartography (Surge AI chart-reading) with tools |
| CursorBench | 43.7% | LowOpus 5.5 Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 18; $1.17/task; 15,811 tokens/task; 28 steps/task |
| CursorBench | 43.7% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; Opus 5.5 cost estimated by Anthropic from Cursor token counts; vendor-reported cost/task $1.18 |
| CursorBench | 52.5% | MediumOpus 5.5 Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 6; $2.91/task; 37,954 tokens/task; 54 steps/task |
| CursorBench | 52.5% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; Opus 5.5 cost estimated by Anthropic from Cursor token counts; vendor-reported cost/task $2.9 |
| CursorBench | 56.0% | HighOpus 5.5 High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 3; $3.97/task; 53,078 tokens/task; 68 steps/task |
| CursorBench | 56.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; Opus 5.5 cost estimated by Anthropic from Cursor token counts; vendor-reported cost/task $3.97 |
| CursorBench | 56.0% | Extra highOpus 5.5 Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 2; $6.98/task; 101,083 tokens/task; 109 steps/task |
| CursorBench | 56.0% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; Opus 5.5 cost estimated by Anthropic from Cursor token counts; vendor-reported cost/task $6.99 |
| CursorBench | 57.8% | MaxOpus 5.5 Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 1; $13.43/task; 218,363 tokens/task; 185 steps/task |
| CursorBench | 57.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; Opus 5.5 cost estimated by Anthropic from Cursor token counts; vendor-reported cost/task $13.43 |
| DeepSWE | 68.4% | MaxOpus 5.5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 74.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | avg of 5 trials unless otherwise noted |
| Design Arena (all categories) | 1399 | Defaultclaude-opus-5-5 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±10.4 SE; 1259 battles; win rate 65.3% |
| Epoch Capabilities Index | 167.3 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 164.0-172.0; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 2050 | Default*claude-opus-5-5 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 7; rubric score 16.79/20; slop 10.43; avg length 6041 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*) |
| FrontierCode | 47.3% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Anthropic: scores decline above medium and mostly recover at max (out-of-scope changes penalized); vendor-reported cost/task $0.4 |
| FrontierCode | 54.6% | Mediumclaude-opus-5-5_medium | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | Claude Code | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5 |
| FrontierCode | 54.6% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Anthropic: scores decline above medium and mostly recover at max (out-of-scope changes penalized); vendor-reported cost/task $0.8 |
| FrontierCode | 54.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Anthropic: scores decline above medium and mostly recover at max (out-of-scope changes penalized); vendor-reported cost/task $1.09 |
| FrontierCode | 51.4% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Anthropic: scores decline above medium and mostly recover at max (out-of-scope changes penalized); vendor-reported cost/task $2.25 |
| FrontierCode | 54.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Anthropic: scores decline above medium and mostly recover at max (out-of-scope changes penalized); vendor-reported cost/task $6.19 |
| FrontierCode v1.1 (Extended) | 65.3% | Mediumadaptive thinking, effort=medium (best) | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | FrontierCode v1.1 Extended; declines above medium, mostly recovers at max |
| FrontierCode v1.1 (Extended) | 63.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code | |
| FrontierMath (Tiers 1-3) | 91.2% | Maxclaude-opus-5-5_max | Independent testIndependentEpoch AI ↗ | 22 Sep 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.7pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 4 | 95.0% | Maxclaude-opus-5-5_max | Independent testIndependentEpoch AI ↗ | 22 Sep 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 4.1pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierSWE | 62.3% | Default | Independent testIndependentFrontierSWE ↗ | 1 Oct 2026 | proximus | FrontierSWE V2, mean@5 over 34 tasks (20h budget); ±9.8; $98.87/trial; 14.2h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort |
| FrontierSWE v2 | 62.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Proximal harness | avg of 5 trials unless otherwise noted |
| GDP.pdf | 25.6% | Loweffort=low | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.76412; competitor number as plotted by OpenAI ("Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback") |
| GDP.pdf | 25.6% | Mediumeffort=medium | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.79532; competitor number as plotted by OpenAI ("Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback") |
| GDP.pdf | 28.8% | Higheffort=high | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.82543; competitor number as plotted by OpenAI ("Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback") |
| GDP.pdf | 26.6% | Extra higheffort=xhigh | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.96427; competitor number as plotted by OpenAI ("Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback") |
| GDP.pdf | 26.2% | Maxeffort=max | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 29 Sep 2026 | — | GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $1.55234; competitor number as plotted by OpenAI ("Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback") |
| GDPval-AA v2.1 | 1224 | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $0.21 |
| GDPval-AA v2.1 | 1576 | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $0.86 |
| GDPval-AA v2.1 | 1692 | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $1.54 |
| GDPval-AA v2.1 | 1820 | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1798.02-1842.19 |
| GDPval-AA v2.1 | 1820 | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $4.21 |
| GDPval-AA v2.1 | 1846 | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1822.93-1869.4 |
| GDPval-AA v2.1 | 1846 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $8.92 |
| Harvey LAB-AA | 91.2% | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| HealthBench Professional | 65.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | avg of 5 trials unless otherwise noted |
| HLE Diamond | 55.0% | Defaultclaude-opus-5-5 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±3.1; entry added 2026-09-03 |
| Humanity's Last Exam | 48.3% | LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-low) |
| Humanity's Last Exam | 54.7% | MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-medium) |
| Humanity's Last Exam | 55.6% | HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-high) |
| Humanity's Last Exam | 57.5% | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-xhigh) |
| Humanity's Last Exam | 61.4% | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5) |
| Humanity's Last Exam | 64.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | avg of 5 trials unless otherwise noted |
| Humanity's Last Exam (with tools) | 67.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | with web search, web fetch, programmatic tool calling, code execution; 1M total token cap |
| IOI (Vals) | 95.1% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.944 stderr; $5.258/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 3.8 | HighClaude Opus 5.5 (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 3/56; Thurstone comparison score (centered at 0); est. win chance 93%; 95% bootstrap 3.676 to 3.882 |
| LMArena Code Arena (WebDev) | 1818 | Maxclaude-opus-5.5-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 1 (CI rank 1-1); 95% CI 1801-1834; 1976 votes |
| LMArena Text - Coding category | 1538 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 11 (CI rank 1-50); 95% CI 1519-1557; 918 votes |
| LMArena Text - Creative Writing | 1515 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 2 (rank range 1-7); 95% CI 1493.9-1536.1; 908 votes; style-controlled |
| LMArena Text - Expert | 1539 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 4 (rank range 1-43); 95% CI 1511.1-1566.8; 448 votes; style-controlled |
| LMArena Text - Hard Prompts | 1533 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 3 (rank range 1-11); 95% CI 1521.1-1545.3; 2437 votes; style-controlled |
| LMArena Text - Instruction Following | 1514 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 2 (rank range 1-10); 95% CI 1498.2-1530.7; 1337 votes; style-controlled |
| LMArena Text - Longer Query | 1528 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 2 (rank range 1-9); 95% CI 1512.9-1542.7; 1708 votes; style-controlled |
| LMArena Text - Multi-Turn | 1497 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 11 (rank range 2-75); 95% CI 1470.6-1523.1; 496 votes; style-controlled |
| LMArena Text - Non-English | 1496 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 2 (rank range 1-16); 95% CI 1484.2-1508.1; 2470 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1493 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 10 (rank range 1-55); 95% CI 1470.7-1515.0; 716 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1490 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 23 (rank range 1-100); 95% CI 1457.7-1523.0; 336 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1519 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 2 (rank range 1-7); 95% CI 1500.8-1537.3; 1159 votes; style-controlled |
| LMArena Text (overall) | 1504 | Highclaude-opus-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 4 (CI rank 2-14); 95% CI 1494-1514; 3932 votes |
| OfficeQA Pro | 67.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | avg of 5 trials unless otherwise noted |
| OSWorld 2.1 (partial score) | 81.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | partial score; system card table labels it OSWorld 2.0 (partial/strict 81.8/48.7) |
| OSWorld 2.1 (strict pass rate) | 48.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | strict pass rate (card labels OSWorld 2.0) |
| PRBench Legal (Scale) | 47.6% | Defaultclaude-opus-5-5 | Independent testIndependentScale AI (SEAL) ↗ | 25 Sep 2026 | — | Scale rank 8; ±2.2 CI |
| ProgramBench (fully resolved) | 18.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±2.753 stderr; $68.051/test; strict fully-resolved rate |
| ProgramBench (fully resolved) | 91.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | mini-swe-agent | avg of 5 trials unless otherwise noted |
| SciCode | 58.6% | LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-low) |
| SciCode | 59.3% | MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-medium) |
| SciCode | 60.4% | HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-high) |
| SciCode | 65.0% | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-xhigh) |
| SciCode | 66.9% | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5) |
| SimpleBench | 88.4% | DefaultClaude Opus 5.5 | Independent testIndependentSimpleBench ↗ | 24 Sep 2026 | — | AVG@5, temp 0.7; rank 1st; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 72.2% | Maxclaude-opus-5-5_max | Independent testIndependentEpoch AI ↗ | 22 Sep 2026 | — | Epoch-run (no tools); ±1.42 stderr |
| SWE Atlas - Codebase QnA | 66.4% | MaxOpus 5.5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE-bench Multilingual | 93.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | avg of 5 trials unless otherwise noted |
| SWE-bench Multimodal | 61.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | avg of 5 trials unless otherwise noted |
| SWE-bench Pro (Anthropic internal subset) | 87.4% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 1 Oct 2026 | — | Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $0.12 per solved task |
| SWE-bench Pro (Anthropic internal subset) | 92.8% | Mediumadaptive thinking, effort=medium (default) | Maker's own figureVendor-reportedAnthropic ↗ | 1 Oct 2026 | — | Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $0.22 per solved task; ~2.5 pts below high for ~70% of the cost |
| SWE-bench Pro (Anthropic internal subset) | 95.3% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 1 Oct 2026 | — | Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $0.29 per task; xhigh adds ~1.4 pts for 2.5x the cost of high |
| SWE-Bench Pro (public, v1) | 89.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | avg of 5 trials unless otherwise noted |
| Terminal-Bench 2.1 | 87.6% | Higheffort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±1.716 stderr; $0.496/test |
| Terminal-Bench 4.0 | 31.3% | LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-low) |
| Terminal-Bench 4.0 | 38.5% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; xhigh is the reported headline (max within noise, SE ±2.6); production safeguards on, fallbacks on 2.5% of requests; vendor-reported cost/task $1.29 |
| Terminal-Bench 4.0 | 52.5% | MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-medium) |
| Terminal-Bench 4.0 | 57.6% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; xhigh is the reported headline (max within noise, SE ±2.6); production safeguards on, fallbacks on 2.5% of requests; vendor-reported cost/task $2.94 |
| Terminal-Bench 4.0 | 56.6% | HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-high) |
| Terminal-Bench 4.0 | 64.2% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; xhigh is the reported headline (max within noise, SE ±2.6); production safeguards on, fallbacks on 2.5% of requests; vendor-reported cost/task $3.88 |
| Terminal-Bench 4.0 | 59.6% | Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5-xhigh) |
| Terminal-Bench 4.0 | 66.4% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; xhigh is the reported headline (max within noise, SE ±2.6); production safeguards on, fallbacks on 2.5% of requests; vendor-reported cost/task $7.35 |
| Terminal-Bench 4.0 | 65.2% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0 stderr; $13.199/test |
| Terminal-Bench 4.0 | 63.1% | MaxOpus 5.5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 59.6% | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-opus-5-5) |
| Terminal-Bench 4.0 | 64.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; xhigh is the reported headline (max within noise, SE ±2.6); production safeguards on, fallbacks on 2.5% of requests; vendor-reported cost/task $11.24 |
| Terminal-Bench-Science 0.1 | 63.3% | Maxeffort=max | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | 29 Sep 2026 | — | Terminal-Bench Science 0.1; vendor-estimated API cost/task $23.2108; competitor number as plotted by OpenAI |
| Terminal-Bench-Science 0.1 | 58.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | Claude Code (--bare) | Terminal-Bench-Science 0.1 (70 tasks); SE ±4.8; safeguards on (fallbacks on 3.9% of requests) |
| Toolathlon | 77.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | avg of 5 trials unless otherwise noted |
| Vals Code Migration | 66.7% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.329 stderr; $112.971/test |
| Vals Index | 67.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±0.888 stderr; $32.144/test |
| Vals Legal Research Bench | 50.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-opus-5-5; rank 5/72; ±3.475 stderr; $25.163725/test |
| Vals Public Benefits Bench v1.1 | 70.6% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-opus-5-5; rank 3/45; ±1.185 stderr; $5.952407/test |
| Vals SRE Bench | 33.6% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±2.923 stderr; $34.654/test |
| Vals Vibe Code Bench | 90.3% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±1.526 stderr; $57.923/test |
| Vending-Bench 2 | $9,235 | Default | Independent testIndependentAndon Labs ↗ | 1 Oct 2026 | — | final money balance after simulated year, arithmetic mean across runs; ±$785; rank 8; only top 10 rendered server-side (57 more behind 'Show more') |
| WANDR | 31.2% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $1.2 |
| WANDR | 62.8% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $11.2 |
| WANDR | 67.3% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $16.88 |
| WANDR | 71.3% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $29.06 |
| WANDR | 72.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 22 Sep 2026 | — | WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $37.92 |