Skip to content
Bencher

Models · Anthropic · Out sinceReleased 30 Jun 2026

Claude Sonnet 5#27 for coding.#27 for coding, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary
Where you can use it:In apps:Claude (Team plan) · Sonnet 5

Claude Sonnet 5 is made by Anthropic. Among the models we track it ranks #27 for coding, #29 for research and analysis, #31 for writing. It's mid-priced to use.

API id claude-sonnet-5. Introductory $2/$10 pricing made permanent Aug 10, 2026. Default effort high. New tokenizer (~1.0-1.35x tokens vs Sonnet 4.6).

Writing & creativity
32.6 / 100 · #31
Research & analysis
55.8 / 100 · #29
Coding
36.0 / 100 · #27
Price
Mid-priced$2 / $10
Price per 1M (blended)Blended / 1M
$4
MemoryContext
1M
Longest answerMax output
128K
Test resultsResults
92 (50 independent50 indep.)
Out sinceReleased
30 Jun 2026
Made byVendor
Anthropic
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra highMax

Highlighted: where it did best for coding (Maximum thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-sonnet-5.

Route it as anthropic/claude-sonnet-5 at $2 in / $10 out per 1M tokens, 1M context. Listed since 30 Jun 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstool_choicetoolsverbosity

Thinking level

Reasoning effort

How long should Claude Sonnet 5Where Claude Sonnet 5think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at MaxBest at Max

CursorBench

Best at MaxBest at Max

FrontierCode

Best at Extra high, worse abovePeaks at Extra high

Terminal-Bench 2.1

Best at MaxBest at Max

SciCode

Best at MaxBest at Max

Artificial Analysis Intelligence Index

Best at MaxBest at Max

AA-Briefcase v1.1

Best at MaxBest at Max

Humanity's Last Exam

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.1923Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $0.82
AA-Briefcase v1.11056Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $1.73
AA-Briefcase v1.11177Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $3.76
AA-Briefcase v1.11274Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $7.56
AA-Briefcase v1.11359Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $14.43
Artificial Analysis Intelligence Index34.4Extra highClaude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-sonnet-5-xhigh; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $2.87/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index38.2MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-sonnet-5; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $5.09/task (AA marks this variant deprecated)
Artificial Analysis output speed63 tok/sExtra highClaude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 20.2s; list price $2/10 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed79 tok/sMaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 212.4s; list price $2/10 per 1M in/out (AA marks this variant deprecated)
AutomationBench10.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
BrowseComp86.6%Maxadaptive thinking, effort=max, multi-agentMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—multi-agent configuration
BrowseComp84.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—single agent; 10M token limit with compaction
Chartography (no tools)15.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—no tools
CursorBench24.1%LowSonnet 5 LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 61; $1.39/task; 23,772 tokens/task; 46 steps/task
CursorBench24.1%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $1.39
CursorBench28.0%MediumSonnet 5 MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 54; $2.31/task; 39,114 tokens/task; 65 steps/task
CursorBench28.0%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $2.31
CursorBench30.8%HighSonnet 5 HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 50; $3.48/task; 61,146 tokens/task; 85 steps/task
CursorBench30.8%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $3.48
CursorBench32.0%Extra highSonnet 5 Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 47; $4.55/task; 83,373 tokens/task; 102 steps/task
CursorBench32.0%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $4.55
CursorBench34.1%MaxSonnet 5 MaxIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 39; $7.17/task; 149,257 tokens/task; 140 steps/task
CursorBench34.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; vendor-reported cost/task $7.17
CursorBench (pre-3.2, legacy)61.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026Cursor agentCursorBench version as of June 2026 (pre-3.2); effort per Cursor-reported best
Design Arena (all categories)1283Defaultclaude-sonnet-5Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.5 SE; 21694 battles; win rate 53.2%
Design Arena (fullstack)1247Defaultclaude-sonnet-5Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±6.3 SE; 3733 battles; win rate 53.7%
EQ-Bench Creative Writing v3 (Elo)1794Defaultclaude-sonnet-5Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 27; rubric score 16.47/20; slop 11.60; avg length 5753 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
Frontier-Bench v0.117.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026mini-SWE-agent (GKE)internal run
FrontierCode28.7%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); vendor-reported cost/task $2.39
FrontierCode35.2%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); vendor-reported cost/task $3.81
FrontierCode39.4%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); vendor-reported cost/task $6.1
FrontierCode42.7%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); vendor-reported cost/task $10.07
FrontierCode42.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); vendor-reported cost/task $17.12
FrontierCode42.7%Defaultclaude-sonnet-5_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026Claude Coderead from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5
FrontierCode v138.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026Claude Code
GDPval-AA v21618Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—Elo as of June 17, 2026
GDPval-AA v2.11449Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations)
GPQA Diamond91.1%MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5) (AA marks this variant deprecated)
GSO37.3%Extra highreasoning_effort=xhighIndependent testIndependentGSO ↗12 Jul 2026OpenHandsOpt@1; hack-controlled score 36.27; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable
HealthBench Professional57.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—
Humanity's Last Exam39.0%Extra highClaude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-xhigh) (AA marks this variant deprecated)
Humanity's Last Exam41.3%MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5) (AA marks this variant deprecated)
Humanity's Last Exam43.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—no tools
Humanity's Last Exam (with tools)54.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
LegalBench (Vals)83.9%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-sonnet-5; rank 44/149; ±0.458 stderr; $0.005759/test
LMArena Search Arena1194Defaultclaude-sonnet-5-searchIndependent testIndependentLMArena ↗24 Aug 2026—rank 17 (rank range 12-18); 95% CI 1187.2-1201.0; 40230 votes
LMArena Text - Creative Writing1436Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 61 (rank range 36-84); 95% CI 1428.4-1443.1; 9284 votes; style-controlled
LMArena Text - Expert1512Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 23 (rank range 8-54); 95% CI 1502.9-1520.6; 5051 votes; style-controlled
LMArena Text - Hard Prompts1491Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 45 (rank range 28-60); 95% CI 1485.8-1495.6; 29825 votes; style-controlled
LMArena Text - Instruction Following1467Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 37 (rank range 20-58); 95% CI 1461.1-1472.7; 15961 votes; style-controlled
LMArena Text - Longer Query1482Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 38 (rank range 21-59); 95% CI 1476.3-1487.4; 21121 votes; style-controlled
LMArena Text - Multi-Turn1474Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 51 (rank range 18-78); 95% CI 1466.0-1481.1; 7677 votes; style-controlled
LMArena Text - Non-English1448Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 55 (rank range 42-77); 95% CI 1443.2-1453.2; 26108 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1465Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 48 (rank range 25-77); 95% CI 1457.7-1472.0; 8766 votes; style-controlled
LMArena Text - Occupational: Legal & Government1478Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 41 (rank range 13-89); 95% CI 1467.6-1487.9; 3876 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1452Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 46 (rank range 29-70); 95% CI 1445.8-1458.8; 12341 votes; style-controlled
OSWorld 2.1 (partial score)57.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—partial score
OSWorld-Verified81.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—updated OSWorld-Verified methodology (zoom-tool bug fix, 128K max tokens/turn)
ProgramBench (fully resolved)0.0%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0 stderr; $24.358/test; strict fully-resolved rate
ProgramBench (fully resolved)77.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
SciCode54.1%Extra highClaude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-xhigh) (AA marks this variant deprecated)
SciCode54.3%MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5) (AA marks this variant deprecated)
SimpleBench60.6%DefaultClaude Sonnet 5Independent testIndependentSimpleBench ↗9 Jul 2026—AVG@5, temp 0.7; rank 34th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SWE-bench Multilingual78.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
SWE-bench Multimodal28.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
SWE-bench Pro (Anthropic internal subset)77.4%Highadaptive thinking, effort=high (default)Maker's own figureVendor-reportedAnthropic ↗1 Oct 2026—Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $0.84 per solved task
SWE-Bench Pro (public, v1)63.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—
SWE-Bench Pro V2 (full)93.2%Extra highSonnet 5 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 8 (Scale rank accounts for CI); ±1.71; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-Bench Pro V2 (hard)88.2%Extra highSonnet 5 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 4 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-bench Verified85.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026—
SWE-rebench56.8%HighSonnet 5 [high]Independent testIndependentSWE-rebench (Nebius) ↗1 Oct 2026SWE-rebench standard scaffoldtime window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±0.94; pass@5 74.8%; $1.43/problem
Terminal-Bench 2.180.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026mini-SWE-agent (GKE)89 tasks x 5 attempts
Terminal-Bench 2.180.5%MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5) (AA marks this variant deprecated)
Terminal-Bench 2.174.5%Maxeffort=maxIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±2.085 stderr; $0.535/test
Terminal-Bench 2.174.6%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Claude CodeTB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 3.014.6%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Claude CodeTB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 17.9B, run cost $6.9k
Terminal-Bench 4.03.2%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; vendor-reported cost/task $3.78
Terminal-Bench 4.04.4%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; vendor-reported cost/task $4.83
Terminal-Bench 4.04.5%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; vendor-reported cost/task $8.2
Terminal-Bench 4.07.1%Extra highClaude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5-xhigh) (AA marks this variant deprecated)
Terminal-Bench 4.07.0%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; vendor-reported cost/task $9.95
Terminal-Bench 4.014.1%MaxClaude Sonnet 5 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-sonnet-5) (AA marks this variant deprecated)
Terminal-Bench 4.012.4%MaxIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.06 95% CI; 330 trials; total run cost $9604; model release 2026-06-30
Terminal-Bench 4.010.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; vendor-reported cost/task $11.62
Toolathlon74.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—internal harness, Pass@1
Vals Code Migration44.4%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.246 stderr; $35.310/test
Vals CorpFin v267.0%Maxeffort=maxIndependent testIndependentVals.ai ↗12 Aug 2026—vals id anthropic/claude-sonnet-5; rank 15/134; ±0.926 stderr; $0.516119/test
Vals Index51.8%Maxeffort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.092 stderr; $13.721/test
Vals Legal Research Bench41.8%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-sonnet-5; rank 20/72; ±3.428 stderr; $2.719446/test
Vals Public Benefits Bench v1.166.0%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-sonnet-5; rank 17/45; ±1.232 stderr; $1.290695/test
Vals TaxEval v275.6%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id anthropic/claude-sonnet-5; rank 15/145; ±0.836 stderr; $0.301619/test
Vals Vibe Code Bench81.3%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±3.046 stderr; $25.385/test