Models · Anthropic · Out sinceReleased 28 Sep 2026
Claude Sonnet 5.5#1 for coding.#1 for coding, best at max effort.
Claude Sonnet 5.5 is made by Anthropic. Among the models we track it ranks #1 for coding, #6 for research and analysis. It's mid-priced to use. Nearly as good as the top pick on coding tests and costs half as much per word.
API id claude-sonnet-5-5. Default effort high on the API (medium in Claude Code and apps). Effort levels recalibrated vs Sonnet 5. thinking:{type:"between_tools"} replaces "disabled" (works at low/medium/high only). Cache reads $0.20/M.
- —
- 79.9 / 100 · #6
- 92.6 / 100 · #1
- Mid-priced$2 / $10
- $4
- 1M
- 128K
- 120 (82 independent82 indep.)
- 28 Sep 2026
- Anthropic
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for coding (Maximum thinking).
Highlighted: dominant setting in its coding composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-sonnet-5.5.
Route it as anthropic/claude-sonnet-5.5 at $2 in / $10 out per 1M tokens, 1M context. Listed since 28 Sep 2026.
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolsverbosity
Thinking level
Reasoning effort
How long should Claude Sonnet 5.5Where Claude Sonnet 5.5think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Artificial Analysis Coding Agent Index
Best at MaxBest at Max
Terminal-Bench 4.0
Best at MaxBest at Max
CursorBench
Best at MaxBest at Max
DeepSWE
Best at MaxBest at Max
FrontierCode
Best at Extra high, worse abovePeaks at Extra high
FrontierCode v1.1 (Extended)
Best at Extra high, worse abovePeaks at Extra high
SWE Atlas - Codebase QnA
Best at MaxBest at Max
SciCode
Best at MaxBest at Max
Artificial Analysis Intelligence Index
Best at MaxBest at Max
AA-Briefcase v1.1
Best at MaxBest at Max
AA-LCR
Best at MaxBest at Max
AA-Omniscience Accuracy
Best at MaxBest at Max
AA-Omniscience Hallucination Rate
Best at MaxBest at Max
Lower is better on this test
AA-Omniscience Index
Best at MaxBest at Max
GDPval-AA v2.1
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1264 | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $0.87 |
| AA-Briefcase v1.1 | 1461 | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $1.64 |
| AA-Briefcase v1.1 | 1634 | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $3.95 |
| AA-Briefcase v1.1 | 1746 | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Briefcase v1.1 | 1746 | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $9.63 |
| AA-Briefcase v1.1 | 1811 | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Briefcase v1.1 | 1811 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $29.19 |
| AA-LCR | 76.3% | MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 78.0% | HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 79.7% | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 82.7% | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 47.2% | MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 52.0% | HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 53.0% | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 54.0% | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 51.2% | MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 64.6% | HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 62.9% | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 47.0% | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 20.1 | MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 21 | HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 23.5 | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 32.3 | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Artificial Analysis Coding Agent Index | 42.1 | LowSonnet 5.5 (low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | agent Claude Code; components: DeepSWE v1.1 61.9, SWE-Atlas-QnA 39.0, Terminal-Bench v4 25.3; avg cost $0.48/task; avg wall time 6 min/task |
| Artificial Analysis Coding Agent Index | 45.9 | MediumSonnet 5.5 (medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | agent Claude Code; components: DeepSWE v1.1 65.5, SWE-Atlas-QnA 44.9, Terminal-Bench v4 27.3; avg cost $0.62/task; avg wall time 9 min/task |
| Artificial Analysis Coding Agent Index | 55 | HighSonnet 5.5 (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | agent Claude Code; components: DeepSWE v1.1 66.7, SWE-Atlas-QnA 56.5, Terminal-Bench v4 41.9; avg cost $1.24/task; avg wall time 12 min/task |
| Artificial Analysis Coding Agent Index | 62.9 | Extra highSonnet 5.5 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | agent Claude Code; components: DeepSWE v1.1 68.4, SWE-Atlas-QnA 62.1, Terminal-Bench v4 58.1; avg cost $3.33/task; avg wall time 27 min/task |
| Artificial Analysis Coding Agent Index | 68.4 | MaxSonnet 5.5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | agent Claude Code; components: DeepSWE v1.1 72.0, SWE-Atlas-QnA 66.9, Terminal-Bench v4 66.2; avg cost $14.19/task; avg wall time 87 min/task |
| Artificial Analysis Intelligence Index | 40.7 | MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-sonnet-5-5-medium; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.59/task |
| Artificial Analysis Intelligence Index | 46.7 | HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-sonnet-5-5-high; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $1.08/task |
| Artificial Analysis Intelligence Index | 51.9 | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-sonnet-5-5-xhigh; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $2.74/task |
| Artificial Analysis Intelligence Index | 56 | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug claude-sonnet-5-5; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $7.62/task |
| Artificial Analysis output speed | 91 tok/s | MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 1.4s; list price $2/10 per 1M in/out |
| Artificial Analysis output speed | 94 tok/s | HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 12.9s; list price $2/10 per 1M in/out |
| Artificial Analysis output speed | 101 tok/s | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 32.5s; list price $2/10 per 1M in/out |
| Artificial Analysis output speed | 139 tok/s | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 438.4s; list price $2/10 per 1M in/out |
| ArXivMath (no tools) | 86.8% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| ArXivMath (with tools) | 95.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| AutomationBench | 44.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| Chartography (no tools) | 61.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | no tools |
| CursorBench | 35.8% | LowSonnet 5.5 Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 37; $0.50/task; 11,668 tokens/task; 18 steps/task |
| CursorBench | 35.8% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $0.5 |
| CursorBench | 39.2% | MediumSonnet 5.5 Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 29; $0.70/task; 16,036 tokens/task; 22 steps/task |
| CursorBench | 39.2% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $0.7 |
| CursorBench | 47.8% | HighSonnet 5.5 High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 10; $1.67/task; 37,391 tokens/task; 41 steps/task |
| CursorBench | 47.8% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $1.67 |
| CursorBench | 53.1% | Extra highSonnet 5.5 Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 5; $3.88/task; 100,158 tokens/task; 78 steps/task |
| CursorBench | 53.1% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $3.88 |
| CursorBench | 55.5% | MaxSonnet 5.5 Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 4; $9.67/task; 271,920 tokens/task; 170 steps/task |
| CursorBench | 55.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Cursor agent | CursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $9.67 |
| DeepSWE | 61.9% | LowSonnet 5.5 (low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 65.5% | MediumSonnet 5.5 (medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 66.7% | HighSonnet 5.5 (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 68.4% | Extra highSonnet 5.5 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 72.0% | MaxSonnet 5.5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 71.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| Epoch Capabilities Index | 165.2 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 161.8-169.4; best of listed model versions |
| FrontierCode | 29.3% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); at max Sonnet 5.5 more often ran Claude Code code-review skill with many subagents -> timeouts/out-of-scope edits, lower score than xhigh; vendor-reported cost/task $0.19 |
| FrontierCode | 36.5% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); at max Sonnet 5.5 more often ran Claude Code code-review skill with many subagents -> timeouts/out-of-scope edits, lower score than xhigh; vendor-reported cost/task $0.24 |
| FrontierCode | 49.4% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); at max Sonnet 5.5 more often ran Claude Code code-review skill with many subagents -> timeouts/out-of-scope edits, lower score than xhigh; vendor-reported cost/task $0.42 |
| FrontierCode | 52.1% | Extra highclaude-sonnet-5-5_xhigh | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | Claude Code | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5 |
| FrontierCode | 52.1% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); at max Sonnet 5.5 more often ran Claude Code code-review skill with many subagents -> timeouts/out-of-scope edits, lower score than xhigh; vendor-reported cost/task $1.59 |
| FrontierCode | 46.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | FrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); at max Sonnet 5.5 more often ran Claude Code code-review skill with many subagents -> timeouts/out-of-scope edits, lower score than xhigh; vendor-reported cost/task $20.78 |
| FrontierCode v1.1 (Extended) | 64.4% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | |
| FrontierCode v1.1 (Extended) | 59.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code | lower at max than xhigh (out-of-scope edits/timeouts) |
| FrontierMath (Tiers 1-3) | 88.8% | Maxclaude-sonnet-5-5_max | Independent testIndependentEpoch AI ↗ | 29 Sep 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath Tier 4 | 80.5% | Maxclaude-sonnet-5-5_max | Independent testIndependentEpoch AI ↗ | 29 Sep 2026 | — | FrontierMath Tier 4 (v2), Epoch-run; stderr 6.3pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv) |
| FrontierSWE v2 | 61.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Proximal harness | |
| GDPval-AA v2.1 | 1725 | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1702.96-1746.84 |
| GDPval-AA v2.1 | 1844 | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1820.57-1867.85 |
| GDPval-AA v2.1 | 1844 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); run on pre-release deployment with a structured-outputs bug (Anthropic expects small understatement) |
| GPQA Diamond | 95.6% | Maxclaude-sonnet-5-5_max | Independent testIndependentEpoch AI ↗ | 29 Sep 2026 | — | Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| Harvey LAB-AA | 93.1% | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| HealthBench Professional | 69.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| Humanity's Last Exam | 39.8% | MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5-medium) |
| Humanity's Last Exam | 45.8% | HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5-high) |
| Humanity's Last Exam | 50.0% | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5-xhigh) |
| Humanity's Last Exam | 55.0% | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5) |
| Humanity's Last Exam | 56.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| Humanity's Last Exam (with tools) | 64.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| IOI (Vals) | 83.1% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±3.737 stderr; $7.524/test |
| LMArena Code Arena (WebDev) | 1709 | Highclaude-sonnet-5.5-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 5 (CI rank 5-7); 95% CI 1693-1724; 1810 votes |
| OSWorld 2.1 (partial score) | 80.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | partial score |
| ProgramBench (fully resolved) | 79.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| SciCode | 52.9% | MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5-medium) |
| SciCode | 53.7% | HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5-high) |
| SciCode | 57.3% | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5-xhigh) |
| SciCode | 61.0% | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5) |
| SimpleQA Verified | 46.5% | Maxclaude-sonnet-5-5_max | Independent testIndependentEpoch AI ↗ | 29 Sep 2026 | — | Epoch-run (no tools); ±1.58 stderr |
| SWE Atlas - Codebase QnA | 39.0% | LowSonnet 5.5 (low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 44.9% | MediumSonnet 5.5 (medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 56.5% | HighSonnet 5.5 (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 62.1% | Extra highSonnet 5.5 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 66.9% | MaxSonnet 5.5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE-bench Multilingual | 90.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| SWE-bench Multimodal | 54.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| SWE-Bench Pro (public, v1) | 81.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | — | |
| Terminal-Bench 2.1 | 83.2% | Higheffort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±1.716 stderr; $0.625/test |
| Terminal-Bench 4.0 | 25.3% | LowSonnet 5.5 (low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 20.0% | Lowadaptive thinking, effort=low | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; no internet egress; SE ±2.5; vendor-reported cost/task $0.76 |
| Terminal-Bench 4.0 | 29.8% | MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5-medium) |
| Terminal-Bench 4.0 | 27.3% | MediumSonnet 5.5 (medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 28.8% | Mediumadaptive thinking, effort=med | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; no internet egress; SE ±2.5; vendor-reported cost/task $0.83 |
| Terminal-Bench 4.0 | 43.9% | HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5-high) |
| Terminal-Bench 4.0 | 41.9% | HighSonnet 5.5 (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 43.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; no internet egress; SE ±2.5; vendor-reported cost/task $1.94 |
| Terminal-Bench 4.0 | 58.1% | Extra highSonnet 5.5 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 57.1% | Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5-xhigh) |
| Terminal-Bench 4.0 | 61.5% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; no internet egress; SE ±2.5; vendor-reported cost/task $5.3 |
| Terminal-Bench 4.0 | 66.2% | MaxSonnet 5.5 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 64.1% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.01 stderr; $16.513/test |
| Terminal-Bench 4.0 | 63.6% | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug claude-sonnet-5-5) |
| Terminal-Bench 4.0 | 70.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; no internet egress; SE ±2.5; vendor-reported cost/task $12.54 |
| Terminal-Bench-Science 0.1 | 59.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | 28 Sep 2026 | Claude Code (--bare) | Terminal-Bench-Science 0.1 (70 tasks); SE ±4.7 |
| Vals Code Migration | 69.8% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.262 stderr; $75.827/test |
| Vals Index | 67.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±0.92 stderr; $21.345/test |
| Vals Legal Research Bench | 48.1% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-sonnet-5-5; rank 10/72; ±3.473 stderr; $14.826571/test |
| Vals Public Benefits Bench v1.1 | 67.2% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id anthropic/claude-sonnet-5-5; rank 13/45; ±1.221 stderr; $4.245484/test |
| Vals SRE Bench | 30.1% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±2.841 stderr; $26.521/test |
| Vals Vibe Code Bench | 92.4% | Maxeffort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±1.26 stderr; $31.251/test |