Skip to content
Bencher

Models · Anthropic · Out sinceReleased 28 May 2026

Claude Opus 4.8#17 for coding.#17 for coding, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Claude Opus 4.8 is made by Anthropic. Among the models we track it ranks #17 for coding, #21 for writing, #28 for research and analysis. It's expensive to use.

API id claude-opus-4-8. Default effort high; Anthropic recommends starting at xhigh for coding/agentic work. Fast mode $10/$50.

Writing & creativity
52.6 / 100 · #21
Research & analysis
57.1 / 100 · #28
Coding
52.3 / 100 · #17
Price
Expensive$5 / $25
Price per 1M (blended)Blended / 1M
$10
MemoryContext
1M
Longest answerMax output
128K
Test resultsResults
102 (67 independent67 indep.)
Out sinceReleased
28 May 2026
Made byVendor
Anthropic
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra highMax

Highlighted: where it did best for coding (Maximum thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-opus-4.8.

Route it as anthropic/claude-opus-4.8 at $5 in / $25 out per 1M tokens, 1M context. Listed since 27 May 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolsverbosity

Thinking level

Reasoning effort

How long should Claude Opus 4.8Where Claude Opus 4.8think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

DeepSWE

Best at MaxBest at Max

FrontierCode

Best at MaxBest at Max

ProgramBench (fully resolved)

Best at MaxBest at Max

Terminal-Bench 3.0

Best at MaxBest at Max

Terminal-Bench 2.1

Best at MaxBest at Max

LLM Creative Story-Writing Benchmark (Lech Mazur)

Best at Extra highBest at Extra high

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase1346Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—
ARC-AGI-192.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—
ARC-AGI-272.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—
ARC-AGI-31.5%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—ARC Prize Foundation; read from launch-post benchmark table image (values printed in table)
Artificial Analysis Intelligence Index41.8MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-4-8; list price $5/25 per 1M in/out; cost to run AA Intelligence Index $4.08/task (AA marks this variant deprecated)
Artificial Analysis output speed59 tok/sMaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 64.1s; list price $5/25 per 1M in/out (AA marks this variant deprecated)
AutomationBench17.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
Blueprint-Bench 214.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026—read from launch-post benchmark table image (values printed in table)
BrowseComp88.5%Maxadaptive thinking, effort=max, multi-agentMaker's own figureVendor-reportedAnthropic ↗28 May 2026—multi-agent
BrowseComp84.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
CursorBench (pre-3.2, legacy)63.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗30 Jun 2026Cursor agentCursorBench version as of June 2026 (pre-3.2); effort per Cursor-reported best
DeepSWE54.4%Extra highclaude-opus-4-8_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 80.5%; ±3.7; 4 runs; $8.01/task
DeepSWE59.0%Maxclaude-opus-4-8_maxIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 79.3%; ±1.8; 4 runs; $13.22/task
DeepSWE59.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
Design Arena (all categories)1259Defaultclaude-opus-4-8Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.2 SE; 29994 battles; win rate 52.8%
Design Arena (fullstack)1247Defaultclaude-opus-4-8Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±7.2 SE; 2698 battles; win rate 55.3%
Epoch Capabilities Index158.3DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 156.1-160.7; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1840Defaultclaude-opus-4-8Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 21; rubric score 16.66/20; slop 13.16; avg length 5842 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
Finance Agent53.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—Finance Agent v2
Frontier-Bench v0.121.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
Frontier-Bench v0.118.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026mini-SWE-agent (GKE)internal run
FrontierCode34.3%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026Claude CodeFrontierCode Main at Fable 5 launch (pass rate 37.3%)
FrontierCode46.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
FrontierCode46.5%Defaultclaude-opus-4-8_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026Claude Coderead from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5
FrontierCode (Diamond)13.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026Claude Coderead from launch-post benchmark table image (values printed in table)
FrontierMath (Tiers 1-3)80.0%Maxclaude-opus-4-8_maxIndependent testIndependentEpoch AI ↗10 Jun 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.4pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 456.1%Maxclaude-opus-4-8_maxIndependent testIndependentEpoch AI ↗10 Jun 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 7.8pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
GDP.pdf22.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026—read from launch-post benchmark table image (values printed in table)
GDPval-AA (v1)1890Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026—read from launch-post benchmark table image (values printed in table)
GDPval-AA v21593Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
GPQA Diamond92.4%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±1.984 stderr; $0.339/test
GPQA Diamond92.0%MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated)
GPQA Diamond93.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—
GSO47.1%Extra highreasoning_effort=xhighIndependent testIndependentGSO ↗12 Jul 2026OpenHandsOpt@1; hack-controlled score 47.06; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable
HealthBench Professional57.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
Humanity's Last Exam48.7%MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated)
Humanity's Last Exam49.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
Humanity's Last Exam (with tools)57.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
LiveCodeBench87.8%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±0.948 stderr; $0.327/test
LLM Creative Story-Writing Benchmark (Lech Mazur)0.3HighClaude Opus 4.8 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 25/56; Thurstone comparison score (centered at 0); est. win chance 54%; 95% bootstrap 0.163 to 0.371; incomplete story set (see README coverage note)
LLM Creative Story-Writing Benchmark (Lech Mazur)0.8Extra highClaude Opus 4.8 (xhigh)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 18/56; Thurstone comparison score (centered at 0); est. win chance 62%; 95% bootstrap 0.666 to 0.888
LMArena Search Arena1204Defaultclaude-opus-4-8Independent testIndependentLMArena ↗24 Aug 2026—rank 12 (rank range 8-17); 95% CI 1197.9-1210.7; 70998 votes
LMArena Text - Coding category1533Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 16 (CI rank 7-31); 95% CI 1527-1539; 17275 votes
LMArena Text - Coding category1530Defaultclaude-opus-4-8Independent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 20 (CI rank 7-40); 95% CI 1524-1536; 17933 votes
LMArena Text - Creative Writing1469Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 16 (rank range 9-34); 95% CI 1462.8-1476.0; 12699 votes; style-controlled
LMArena Text - Expert1525Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 14 (rank range 4-35); 95% CI 1517.4-1533.1; 7004 votes; style-controlled
LMArena Text - Expert1517Defaultclaude-opus-4-8Independent testIndependentLMArena ↗30 Sep 2026—rank 18 (rank range 8-46); 95% CI 1508.8-1524.4; 7328 votes; style-controlled
LMArena Text - Hard Prompts1513Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 14 (rank range 7-24); 95% CI 1508.3-1517.4; 42198 votes; style-controlled
LMArena Text - Instruction Following1490Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 12 (rank range 6-21); 95% CI 1485.1-1495.8; 22488 votes; style-controlled
LMArena Text - Instruction Following1479Defaultclaude-opus-4-8Independent testIndependentLMArena ↗30 Sep 2026—rank 19 (rank range 13-39); 95% CI 1473.2-1483.9; 23069 votes; style-controlled
LMArena Text - Longer Query1502Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 12 (rank range 7-26); 95% CI 1496.9-1507.2; 30035 votes; style-controlled
LMArena Text - Longer Query1500Defaultclaude-opus-4-8Independent testIndependentLMArena ↗30 Sep 2026—rank 15 (rank range 7-28); 95% CI 1494.6-1504.9; 30720 votes; style-controlled
LMArena Text - Multi-Turn1498Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 10 (rank range 7-33); 95% CI 1491.0-1504.4; 11116 votes; style-controlled
LMArena Text - Multi-Turn1495Defaultclaude-opus-4-8Independent testIndependentLMArena ↗30 Sep 2026—rank 17 (rank range 7-39); 95% CI 1487.8-1501.3; 11326 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1489Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 12 (rank range 3-32); 95% CI 1482.9-1495.7; 12641 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1487Defaultclaude-opus-4-8Independent testIndependentLMArena ↗30 Sep 2026—rank 16 (rank range 6-36); 95% CI 1480.3-1493.1; 12693 votes; style-controlled
LMArena Text - Occupational: Legal & Government1495Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 16 (rank range 3-50); 95% CI 1486.4-1504.2; 5385 votes; style-controlled
LMArena Text - Occupational: Legal & Government1497Defaultclaude-opus-4-8Independent testIndependentLMArena ↗30 Sep 2026—rank 15 (rank range 2-46); 95% CI 1488.0-1505.7; 5472 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1475Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 18 (rank range 9-36); 95% CI 1468.7-1480.4; 17150 votes; style-controlled
LMArena Text (overall)1481Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 23 (CI rank 13-40); 95% CI 1477-1485; 63577 votes
MCP Atlas82.2%Maxclaude-opus-4-8 (max)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±2.4; entry added 2026-05-28
MCP Atlas82.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—
OSWorld 2.020.6%Maxclaude-opus-4-8_maxIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.5479999999999999; tool setting batched tool; step budget 500
OSWorld 2.055.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—read from launch-post benchmark table image (values printed in table)
OSWorld-Verified83.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026—read from launch-post benchmark table image (values printed in table)
ProgramBench (avg test pass rate)70.9%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 16.5%; avg cost $21.02/task
ProgramBench (fully resolved)0.0%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 16.5%; avg cost $21.02/task
ProgramBench (fully resolved)1.0%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.705 stderr; $31.267/test; strict fully-resolved rate
SciCode54.4%MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated)
ScreenSpot-Pro82.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—no tools
SimpleBench64.8%DefaultClaude Opus 4.8Independent testIndependentSimpleBench ↗29 May 2026—AVG@5, temp 0.7; rank 23rd; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified53.0%Maxclaude-opus-4-8_maxIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
SkillsBench59.2%Maxeffort=maxIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.5 stderr; $4.767/test
SWE Atlas - Codebase QnA57.3%Extra highOpus 4.8 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±4.93; entry added 2026-06-08
SWE Atlas - Refactoring46.7%DefaultOpus 4.8 (Claude Code)Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±6.75; entry added 2026-06-08
SWE Atlas - Test Writing49.6%Extra highOpus 4.8 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±5.93; entry added 2026-06-08
SWE-bench Multilingual84.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—
SWE-bench Multimodal38.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—
SWE-Bench Pro (public, v1)69.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗24 Jul 2026—avg of 5 trials
SWE-bench Verified88.6%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.423 stderr; $1.923/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
SWE-bench Verified88.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—
SWE-bench Verified85.8%DefaultIndependent testIndependentVals.ai ↗1 Sep 2026Claude Code±1.563 stderr; $0.672/test; Vals variant run through Claude Code harness; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
tau2-bench94.4%MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated)
Terminal-Bench 2.182.7%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗9 Jun 2026mini-SWE-agent (GKE)
Terminal-Bench 2.174.6%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗28 May 2026Terminus-2 (Harbor, Daytona)89 tasks x 5 attempts
Terminal-Bench 2.184.6%MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated)
Terminal-Bench 2.171.9%Maxeffort=maxIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.649 stderr; $2.410/test
Terminal-Bench 2.178.9%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Claude CodeTB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 3.015.0%Highshipped default effort (high)Maker's own figureVendor-reportedAnthropic ↗1 Oct 2026Claude Managed Agents74 tasks, two runs per model; $63 per solved task
Terminal-Bench 3.021.1%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Claude CodeTB 3.0 v0.1 (74 tasks); superseded by TB 4.0; tokens 5.2B, run cost $5.2k
Terminal-Bench 4.023.6%MaxIndependent testIndependentTerminal-Bench ↗1 Oct 2026Claude CodeTB 4.0.0 (66 tasks); ±3.56 95% CI; 330 trials; total run cost $6481; model release 2026-05-28
Terminal-Bench 4.023.2%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.336 stderr; $17.141/test
Terminal-Bench 4.021.7%MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-4-8) (AA marks this variant deprecated)
Toolathlon79.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—internal harness, Pass@1
Vals Code Migration47.3%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.176 stderr; $30.514/test
Vals CorpFin v266.7%Maxeffort=maxIndependent testIndependentVals.ai ↗12 Aug 2026—vals id anthropic/claude-opus-4-8; rank 18/134; ±0.928 stderr; $0.855534/test
Vals Index55.1%Maxeffort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.009 stderr; $13.136/test
Vals Legal Research Bench43.8%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-opus-4-8; rank 19/72; ±3.448 stderr; $2.820942/test
Vals Public Benefits Bench v1.168.1%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-opus-4-8; rank 11/45; ±1.212 stderr; $1.889795/test
Vals TaxEval v275.6%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id anthropic/claude-opus-4-8; rank 14/145; ±0.838 stderr; $0.159139/test
Vals Vibe Code Bench82.7%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±3.081 stderr; $26.880/test
Vals Vibe Code Bench77.5%DefaultIndependent testIndependentVals.ai ↗29 Sep 2026Claude Code±3.737 stderr; $6.856/test; Vals variant run through Claude Code harness