Skip to content
Bencher

Models · Anthropic · Out sinceReleased 28 Sep 2026

Claude Sonnet 5.5#1 for coding.#1 for coding, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)ProprietaryOur pick: Great code for half the pricePick: Best coding value
Where you can use it:In apps:Claude (Team plan) · Sonnet 5.5

Claude Sonnet 5.5 is made by Anthropic. Among the models we track it ranks #1 for coding, #6 for research and analysis. It's mid-priced to use. Nearly as good as the top pick on coding tests and costs half as much per word.

API id claude-sonnet-5-5. Default effort high on the API (medium in Claude Code and apps). Effort levels recalibrated vs Sonnet 5. thinking:{type:"between_tools"} replaces "disabled" (works at low/medium/high only). Cache reads $0.20/M.

Writing & creativity
—
Research & analysis
79.9 / 100 · #6
Coding
92.6 / 100 · #1
Price
Mid-priced$2 / $10
Price per 1M (blended)Blended / 1M
$4
MemoryContext
1M
Longest answerMax output
128K
Test resultsResults
120 (82 independent82 indep.)
Out sinceReleased
28 Sep 2026
Made byVendor
Anthropic
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra highMax

Highlighted: where it did best for coding (Maximum thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-sonnet-5.5.

Route it as anthropic/claude-sonnet-5.5 at $2 in / $10 out per 1M tokens, 1M context. Listed since 28 Sep 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolsverbosity

Thinking level

Reasoning effort

How long should Claude Sonnet 5.5Where Claude Sonnet 5.5think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Artificial Analysis Coding Agent Index

Best at MaxBest at Max

Terminal-Bench 4.0

Best at MaxBest at Max

CursorBench

Best at MaxBest at Max

DeepSWE

Best at MaxBest at Max

FrontierCode

Best at Extra high, worse abovePeaks at Extra high

FrontierCode v1.1 (Extended)

Best at Extra high, worse abovePeaks at Extra high

SWE Atlas - Codebase QnA

Best at MaxBest at Max

SciCode

Best at MaxBest at Max

Artificial Analysis Intelligence Index

Best at MaxBest at Max

AA-Briefcase v1.1

Best at MaxBest at Max

AA-LCR

Best at MaxBest at Max

AA-Omniscience Accuracy

Best at MaxBest at Max

AA-Omniscience Hallucination Rate

Best at MaxBest at Max

Lower is better on this test

AA-Omniscience Index

Best at MaxBest at Max

GDPval-AA v2.1

Best at MaxBest at Max

Humanity's Last Exam

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.11264Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $0.87
AA-Briefcase v1.11461Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $1.64
AA-Briefcase v1.11634Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $3.95
AA-Briefcase v1.11746Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Briefcase v1.11746Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $9.63
AA-Briefcase v1.11811MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Briefcase v1.11811Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $29.19
AA-LCR76.3%MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR78.0%HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR79.7%Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR82.7%MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy47.2%MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy52.0%HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy53.0%Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy54.0%MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate51.2%MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate64.6%HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate62.9%Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate47.0%MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index20.1MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index21HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index23.5Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index32.3MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Artificial Analysis Coding Agent Index42.1LowSonnet 5.5 (low)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude Codeagent Claude Code; components: DeepSWE v1.1 61.9, SWE-Atlas-QnA 39.0, Terminal-Bench v4 25.3; avg cost $0.48/task; avg wall time 6 min/task
Artificial Analysis Coding Agent Index45.9MediumSonnet 5.5 (medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude Codeagent Claude Code; components: DeepSWE v1.1 65.5, SWE-Atlas-QnA 44.9, Terminal-Bench v4 27.3; avg cost $0.62/task; avg wall time 9 min/task
Artificial Analysis Coding Agent Index55HighSonnet 5.5 (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude Codeagent Claude Code; components: DeepSWE v1.1 66.7, SWE-Atlas-QnA 56.5, Terminal-Bench v4 41.9; avg cost $1.24/task; avg wall time 12 min/task
Artificial Analysis Coding Agent Index62.9Extra highSonnet 5.5 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude Codeagent Claude Code; components: DeepSWE v1.1 68.4, SWE-Atlas-QnA 62.1, Terminal-Bench v4 58.1; avg cost $3.33/task; avg wall time 27 min/task
Artificial Analysis Coding Agent Index68.4MaxSonnet 5.5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude Codeagent Claude Code; components: DeepSWE v1.1 72.0, SWE-Atlas-QnA 66.9, Terminal-Bench v4 66.2; avg cost $14.19/task; avg wall time 87 min/task
Artificial Analysis Intelligence Index40.7MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-sonnet-5-5-medium; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $0.59/task
Artificial Analysis Intelligence Index46.7HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-sonnet-5-5-high; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $1.08/task
Artificial Analysis Intelligence Index51.9Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-sonnet-5-5-xhigh; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $2.74/task
Artificial Analysis Intelligence Index56MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-sonnet-5-5; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $7.62/task
Artificial Analysis output speed91 tok/sMediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 1.4s; list price $2/10 per 1M in/out
Artificial Analysis output speed94 tok/sHighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 12.9s; list price $2/10 per 1M in/out
Artificial Analysis output speed101 tok/sExtra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 32.5s; list price $2/10 per 1M in/out
Artificial Analysis output speed139 tok/sMaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 438.4s; list price $2/10 per 1M in/out
ArXivMath (no tools)86.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
ArXivMath (with tools)95.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
AutomationBench44.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
Chartography (no tools)61.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—no tools
CursorBench35.8%LowSonnet 5.5 LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 37; $0.50/task; 11,668 tokens/task; 18 steps/task
CursorBench35.8%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $0.5
CursorBench39.2%MediumSonnet 5.5 MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 29; $0.70/task; 16,036 tokens/task; 22 steps/task
CursorBench39.2%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $0.7
CursorBench47.8%HighSonnet 5.5 HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 10; $1.67/task; 37,391 tokens/task; 41 steps/task
CursorBench47.8%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $1.67
CursorBench53.1%Extra highSonnet 5.5 Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 5; $3.88/task; 100,158 tokens/task; 78 steps/task
CursorBench53.1%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $3.88
CursorBench55.5%MaxSonnet 5.5 MaxIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 4; $9.67/task; 271,920 tokens/task; 170 steps/task
CursorBench55.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $9.67
DeepSWE61.9%LowSonnet 5.5 (low)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE65.5%MediumSonnet 5.5 (medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE66.7%HighSonnet 5.5 (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE68.4%Extra highSonnet 5.5 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE72.0%MaxSonnet 5.5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE71.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
Epoch Capabilities Index165.2DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 161.8-169.4; best of listed model versions
FrontierCode29.3%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); at max Sonnet 5.5 more often ran Claude Code code-review skill with many subagents -> timeouts/out-of-scope edits, lower score than xhigh; vendor-reported cost/task $0.19
FrontierCode36.5%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); at max Sonnet 5.5 more often ran Claude Code code-review skill with many subagents -> timeouts/out-of-scope edits, lower score than xhigh; vendor-reported cost/task $0.24
FrontierCode49.4%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); at max Sonnet 5.5 more often ran Claude Code code-review skill with many subagents -> timeouts/out-of-scope edits, lower score than xhigh; vendor-reported cost/task $0.42
FrontierCode52.1%Extra highclaude-sonnet-5-5_xhighIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026Claude Coderead from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5
FrontierCode52.1%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); at max Sonnet 5.5 more often ran Claude Code code-review skill with many subagents -> timeouts/out-of-scope edits, lower score than xhigh; vendor-reported cost/task $1.59
FrontierCode46.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); at max Sonnet 5.5 more often ran Claude Code code-review skill with many subagents -> timeouts/out-of-scope edits, lower score than xhigh; vendor-reported cost/task $20.78
FrontierCode v1.1 (Extended)64.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code
FrontierCode v1.1 (Extended)59.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Codelower at max than xhigh (out-of-scope edits/timeouts)
FrontierMath (Tiers 1-3)88.8%Maxclaude-sonnet-5-5_maxIndependent testIndependentEpoch AI ↗29 Sep 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 480.5%Maxclaude-sonnet-5-5_maxIndependent testIndependentEpoch AI ↗29 Sep 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 6.3pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierSWE v261.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Proximal harness
GDPval-AA v2.11725Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1702.96-1746.84
GDPval-AA v2.11844MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1820.57-1867.85
GDPval-AA v2.11844Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); run on pre-release deployment with a structured-outputs bug (Anthropic expects small understatement)
GPQA Diamond95.6%Maxclaude-sonnet-5-5_maxIndependent testIndependentEpoch AI ↗29 Sep 2026—Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
Harvey LAB-AA93.1%MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
HealthBench Professional69.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
Humanity's Last Exam39.8%MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5-medium)
Humanity's Last Exam45.8%HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5-high)
Humanity's Last Exam50.0%Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5-xhigh)
Humanity's Last Exam55.0%MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5)
Humanity's Last Exam56.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
Humanity's Last Exam (with tools)64.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
IOI (Vals)83.1%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±3.737 stderr; $7.524/test
LMArena Code Arena (WebDev)1709Highclaude-sonnet-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 5 (CI rank 5-7); 95% CI 1693-1724; 1810 votes
OSWorld 2.1 (partial score)80.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—partial score
ProgramBench (fully resolved)79.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
SciCode52.9%MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5-medium)
SciCode53.7%HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5-high)
SciCode57.3%Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5-xhigh)
SciCode61.0%MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5)
SimpleQA Verified46.5%Maxclaude-sonnet-5-5_maxIndependent testIndependentEpoch AI ↗29 Sep 2026—Epoch-run (no tools); ±1.58 stderr
SWE Atlas - Codebase QnA39.0%LowSonnet 5.5 (low)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA44.9%MediumSonnet 5.5 (medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA56.5%HighSonnet 5.5 (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA62.1%Extra highSonnet 5.5 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA66.9%MaxSonnet 5.5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE-bench Multilingual90.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
SWE-bench Multimodal54.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
SWE-Bench Pro (public, v1)81.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
Terminal-Bench 2.183.2%Higheffort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±1.716 stderr; $0.625/test
Terminal-Bench 4.025.3%LowSonnet 5.5 (low)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.020.0%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; no internet egress; SE ±2.5; vendor-reported cost/task $0.76
Terminal-Bench 4.029.8%MediumClaude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5-medium)
Terminal-Bench 4.027.3%MediumSonnet 5.5 (medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.028.8%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; no internet egress; SE ±2.5; vendor-reported cost/task $0.83
Terminal-Bench 4.043.9%HighClaude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5-high)
Terminal-Bench 4.041.9%HighSonnet 5.5 (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.043.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; no internet egress; SE ±2.5; vendor-reported cost/task $1.94
Terminal-Bench 4.058.1%Extra highSonnet 5.5 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.057.1%Extra highClaude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5-xhigh)
Terminal-Bench 4.061.5%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; no internet egress; SE ±2.5; vendor-reported cost/task $5.3
Terminal-Bench 4.066.2%MaxSonnet 5.5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.064.1%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.01 stderr; $16.513/test
Terminal-Bench 4.063.6%MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-5)
Terminal-Bench 4.070.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; no internet egress; SE ±2.5; vendor-reported cost/task $12.54
Terminal-Bench-Science 0.159.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); SE ±4.7
Vals Code Migration69.8%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.262 stderr; $75.827/test
Vals Index67.0%Maxeffort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±0.92 stderr; $21.345/test
Vals Legal Research Bench48.1%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-sonnet-5-5; rank 10/72; ±3.473 stderr; $14.826571/test
Vals Public Benefits Bench v1.167.2%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-sonnet-5-5; rank 13/45; ±1.221 stderr; $4.245484/test
Vals SRE Bench30.1%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±2.841 stderr; $26.521/test
Vals Vibe Code Bench92.4%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±1.26 stderr; $31.251/test