Skip to content
Bencher

Models · Anthropic · Out sinceReleased 16 Apr 2026

Claude Opus 4.7#7 for writing.#7 for writing, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Claude Opus 4.7 is made by Anthropic. Among the models we track it ranks #7 for writing, #25 for coding. It's expensive to use.

API id claude-opus-4-7. Context/max output from OpenRouter listing. Included as older reference point.

Writing & creativity
69.5 / 100 · #7
Research & analysis
—
Coding
41.0 / 100 · #25
Price
Expensive$5 / $25
Price per 1M (blended)Blended / 1M
$10
MemoryContext
1M
Longest answerMax output
128K
Test resultsResults
68 (55 independent55 indep.)
Out sinceReleased
16 Apr 2026
Made byVendor
Anthropic
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra highMax

Highlighted: where it did best for writing (High thinking).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-opus-4.7.

Route it as anthropic/claude-opus-4.7 at $5 in / $25 out per 1M tokens, 1M context. Listed since 16 Apr 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstool_choicetoolsverbosity

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Artificial Analysis Intelligence Index40.7MaxClaude Opus 4.7 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-4-7; list price $5/25 per 1M in/out (AA marks this variant deprecated); estimated
Artificial Analysis output speed46 tok/sMaxClaude Opus 4.7 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 25.2s; list price $5/25 per 1M in/out (AA marks this variant deprecated)
BrowseComp79.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—as listed in Opus 4.8 system card
Epoch Capabilities Index156.3DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 154.4-158.4; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1914Defaultclaude-opus-4-7Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 14; rubric score 16.57/20; slop 11.09; avg length 5692 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
Finance Agent51.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—Finance Agent v2
FrontierCode38.5%Defaultclaude-opus-4-7_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026Claude Coderead from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5
FrontierMath (Tiers 1-3)70.2%Maxclaude-opus-4-7_maxIndependent testIndependentEpoch AI ↗10 Jun 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.7pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
GDPval-AA (v1)1753Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—
GPQA Diamond91.4%MaxClaude Opus 4.7 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-4-7) (AA marks this variant deprecated)
GPQA Diamond94.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—as listed in Opus 4.8 system card
GSO44.1%Highreasoning_effort=highIndependent testIndependentGSO ↗27 Apr 2026OpenHandsOpt@1; hack-controlled score 42.16; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable
Humanity's Last Exam42.3%MaxClaude Opus 4.7 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-4-7) (AA marks this variant deprecated)
Humanity's Last Exam46.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—no tools
Humanity's Last Exam36.2%Defaultclaude-opus-4-7Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 6 (Scale rank accounts for CI); ±1.88; entry added 2026-04-22
Humanity's Last Exam (with tools)54.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—
LLM Creative Story-Writing Benchmark (Lech Mazur)1.8HighClaude Opus 4.7 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 14/56; Thurstone comparison score (centered at 0); est. win chance 75%; 95% bootstrap 1.683 to 1.872; incomplete story set (see README coverage note)
LMArena Search Arena1233Defaultclaude-opus-4-7Independent testIndependentLMArena ↗24 Aug 2026—rank 4 (rank range 3-6); 95% CI 1227.9-1238.6; 91394 votes
LMArena Text - Coding category1551Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 4 (CI rank 1-13); 95% CI 1545-1557; 18437 votes
LMArena Text - Coding category1547Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 6 (CI rank 1-17); 95% CI 1541-1552; 18642 votes
LMArena Text - Creative Writing1489Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 6 (rank range 2-12); 95% CI 1482.3-1496.3; 11712 votes; style-controlled
LMArena Text - Creative Writing1483Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—rank 9 (rank range 5-20); 95% CI 1476.5-1490.3; 11940 votes; style-controlled
LMArena Text - Expert1534Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 8 (rank range 1-23); 95% CI 1525.4-1541.7; 6848 votes; style-controlled
LMArena Text - Expert1534Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—rank 7 (rank range 1-23); 95% CI 1525.9-1542.2; 7015 votes; style-controlled
LMArena Text - Hard Prompts1526Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 6 (rank range 2-10); 95% CI 1521.2-1530.2; 43351 votes; style-controlled
LMArena Text - Hard Prompts1520Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—rank 8 (rank range 4-19); 95% CI 1515.0-1524.0; 44062 votes; style-controlled
LMArena Text - Instruction Following1503Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 5 (rank range 3-10); 95% CI 1497.0-1507.9; 22532 votes; style-controlled
LMArena Text - Instruction Following1493Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—rank 9 (rank range 4-20); 95% CI 1487.6-1498.3; 23111 votes; style-controlled
LMArena Text - Longer Query1515Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 6 (rank range 2-11); 95% CI 1509.7-1520.1; 29893 votes; style-controlled
LMArena Text - Longer Query1508Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—rank 8 (rank range 4-21); 95% CI 1502.8-1513.2; 30621 votes; style-controlled
LMArena Text - Multi-Turn1516Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 5 (rank range 2-10); 95% CI 1509.2-1523.0; 11377 votes; style-controlled
LMArena Text - Multi-Turn1514Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—rank 6 (rank range 2-10); 95% CI 1507.6-1521.3; 11645 votes; style-controlled
LMArena Text - Non-English1490Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 6 (rank range 2-15); 95% CI 1485.6-1494.9; 34660 votes; style-controlled
LMArena Text - Non-English1483Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—rank 13 (rank range 3-21); 95% CI 1478.6-1487.9; 35422 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1508Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 3 (rank range 1-8); 95% CI 1501.8-1514.9; 12918 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1495Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—rank 8 (rank range 3-26); 95% CI 1488.6-1501.6; 13136 votes; style-controlled
LMArena Text - Occupational: Legal & Government1512Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 3 (rank range 1-24); 95% CI 1503.3-1521.5; 5217 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1497Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 6 (rank range 2-9); 95% CI 1491.1-1503.3; 16440 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1484Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—rank 10 (rank range 6-24); 95% CI 1478.1-1490.2; 16859 votes; style-controlled
LMArena Text (overall)1502Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 5 (CI rank 2-11); 95% CI 1498-1505; 64607 votes
LMArena Text (overall)1494Defaultclaude-opus-4-7Independent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 9 (CI rank 4-17); 95% CI 1491-1498; 65705 votes
MCP Atlas79.1%Maxclaude-opus-4-7 (max)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±2.5; entry added 2026-04-08
MCP Atlas79.1%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—as listed in Opus 4.8 system card
OSWorld 2.018.2%Maxclaude-opus-4-7_maxIndependent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.4891; tool setting batched tool; step budget 500
OSWorld-Verified82.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—launch table value; footnote says re-run score 82.3%
ProgramBench (avg test pass rate)55.1%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 4.5%; avg cost $10.96/task
ProgramBench (avg test pass rate)50.9%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 3.0%; avg cost $3.81/task
ProgramBench (fully resolved)0.0%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 4.5%; avg cost $10.96/task
ProgramBench (fully resolved)0.0%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 3.0%; avg cost $3.81/task
SimpleBench61.7%DefaultClaude Opus 4.7Independent testIndependentSimpleBench ↗22 Apr 2026—AVG@5, temp 0.7; rank 29th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified51.7%Extra highclaude-opus-4-7_xhighIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
SWE Atlas - Codebase QnA40.3%DefaultOpus 4.7 (Claude Code)Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 5 (Scale rank accounts for CI); ±5.06; entry added 2026-06-18
SWE Atlas - Refactoring48.6%DefaultOpus-4.7 (Claude Code)Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 1 (Scale rank accounts for CI); ±6.73; entry added 2026-05-06
SWE Atlas - Test Writing38.5%DefaultOpus 4.7 (Claude Code)Independent testIndependentScale AI SEAL ↗1 Oct 2026Claude Coderank 2 (Scale rank accounts for CI); ±5.93; entry added 2026-06-18
SWE-bench Multilingual80.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—as listed in Opus 4.8 system card
SWE-bench Multimodal34.5%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—as listed in Opus 4.8 system card
SWE-Bench Pro (public, v1)64.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 May 2026—
SWE-bench Verified83.5%Maxclaude-opus-4-7_maxIndependent testIndependentEpoch AI ↗20 Apr 2026—Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.7pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv)
SWE-bench Verified82.0%Maxeffort=maxIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.72 stderr; $2.423/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
tau2-bench88.6%MaxClaude Opus 4.7 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-4-7) (AA marks this variant deprecated)
Terminal-Bench 2.183.1%MaxClaude Opus 4.7 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-4-7) (AA marks this variant deprecated)
Terminal-Bench 2.168.9%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Claude CodeTB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.166.1%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Terminus 2TB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.166.1%Defaultas listed (Terminus-2)Maker's own figureVendor-reportedAnthropic ↗28 May 2026Terminus-2 (Harbor)
Terminal-Bench 3.07.0%Highshipped default effort (high)Maker's own figureVendor-reportedAnthropic ↗1 Oct 2026Claude Managed Agents74 tasks, two runs per model; $183 per solved task
Vals Code Migration43.9%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.22 stderr; $36.388/test
Vectara Hallucination Leaderboard (HHEM)12.0%Defaultanthropic/claude-opus-4-7Independent testIndependentVectara ↗22 Sep 2026—factual consistency 88.0 %; answer rate 98.0 %; avg summary 149.1 words; HHEM-2.3 judge; effort not stated (API default)
Vending-Bench 2$10,937DefaultIndependent testIndependentAndon Labs ↗1 Oct 2026—final money balance after simulated year, arithmetic mean across runs; ±$1,181; rank 5; only top 10 rendered server-side (57 more behind 'Show more')