Skip to content
Bencher

Models · Anthropic · Out sinceReleased 22 Sep 2026

Claude Opus 5.5#2 for writing.#2 for writing, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)ProprietaryOur pick: The best writer you can use todayPick: Best for writing & creativityOur pick: The best analyst for long, demanding documentsPick: Best for research & analysisOur pick: The best model for coding right nowPick: Best for coding right now
Where you can use it:In apps:Claude (Team plan) · Opus 5.5

Claude Opus 5.5 is made by Anthropic. Among the models we track it ranks #2 for writing, #2 for coding, #4 for research and analysis. It's expensive to use. In blind tests where thousands of people compare two anonymous texts, people prefer Claude Opus 5.5's writing over every other model you can use today. It's also the strongest at long, well-built stories and texts.

API id claude-opus-5-5. Default effort = medium (unlike other Claude models, which default to high). Adaptive thinking always on (thinking cannot be disabled). Cache reads $0.20/M, cache writes $5/M. Fast mode (up to 2.5x speed) $8/$40. Cyber/bio safeguards with fallback to Opus 4.8 / Opus 5. Knowledge cutoff Jun 2026.

Writing & creativity
86.6 / 100 · #2
Research & analysis
81.2 / 100 · #4
Coding
89.9 / 100 · #2
Price
Expensive$4 / $20
Price per 1M (blended)Blended / 1M
$8
MemoryContext
1M
Longest answerMax output
128K
Test resultsResults
168 (110 independent110 indep.)
Out sinceReleased
22 Sep 2026
Made byVendor
Anthropic
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHighExtra highMax

Highlighted: where it did best for writing (High thinking).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as anthropic/claude-opus-5.5.

Route it as anthropic/claude-opus-5.5 at $4 in / $20 out per 1M tokens, 1M context. Listed since 22 Sep 2026.

include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstemperaturetool_choicetoolsverbosity

Thinking level

Reasoning effort

How long should Claude Opus 5.5Where Claude Opus 5.5think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at MaxBest at Max

CursorBench

Best at MaxBest at Max

FrontierCode

Best at Medium, worse abovePeaks at Medium

FrontierCode v1.1 (Extended)

Best at Medium, worse abovePeaks at Medium

SciCode

Best at MaxBest at Max

SWE-bench Pro (Anthropic internal subset)

Best at HighBest at High

Artificial Analysis Intelligence Index

Best at MaxBest at Max

AutomationBench

Best at MaxBest at Max

AA-Briefcase v1.1

Best at MaxBest at Max

AA-LCR

Best at Extra highBest at Extra high

AA-Omniscience Accuracy

Best at MaxBest at Max

AA-Omniscience Hallucination Rate

Best at MaxBest at Max

Lower is better on this test

AA-Omniscience Index

Best at MaxBest at Max

ARC-AGI-2

Best at High, worse abovePeaks at High

GDP.pdf

Best at High, worse abovePeaks at High

GDPval-AA v2.1

Best at MaxBest at Max

Humanity's Last Exam

Best at MaxBest at Max

WANDR

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.11285Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $1.15
AA-Briefcase v1.11642Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $4.4
AA-Briefcase v1.11705HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Briefcase v1.11705Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $6.27
AA-Briefcase v1.11780Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Briefcase v1.11780Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $12.27
AA-Briefcase v1.11822MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Briefcase v1.11822Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—AA-Briefcase v1.1 (Artificial Analysis, long-horizon knowledge work, Elo); vendor-reported cost/task $21.05
AA-LCR80.7%LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR84.3%MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR82.7%HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR84.7%Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR84.7%MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy63.5%LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy64.5%MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy64.6%HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy65.4%Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy66.2%MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate67.6%LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate68.4%MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate67.6%HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate65.7%Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate58.6%MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index38.9LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index40.3MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index40.6HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index42.7Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index46.4MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
ARC-AGI-287.5%MediumClaude Opus 5.5 (Medium)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.34/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-293.3%HighClaude Opus 5.5 (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.41/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-292.5%Extra highClaude Opus 5.5 (XHigh)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.67/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-291.7%MaxClaude Opus 5.5 (Max)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $1.85/task; data https://arcprize.org/media/data/leaderboard/v2.json
Artificial Analysis Coding Agent Index66MaxOpus 5.5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude Codeagent Claude Code; components: DeepSWE v1.1 68.4, SWE-Atlas-QnA 66.4, Terminal-Bench v4 63.1; avg cost $13.04/task; avg wall time 64 min/task
Artificial Analysis Intelligence Index42.3LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-5-5-low; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $0.55/task
Artificial Analysis Intelligence Index51.2MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-5-5-medium; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $1.34/task
Artificial Analysis Intelligence Index53.6HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-5-5-high; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $1.82/task
Artificial Analysis Intelligence Index56Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-5-5-xhigh; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $3.46/task
Artificial Analysis Intelligence Index57.6MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug claude-opus-5-5; list price $4/20 per 1M in/out; cost to run AA Intelligence Index $5.98/task
Artificial Analysis output speed72 tok/sLowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 13.5s; list price $4/20 per 1M in/out
Artificial Analysis output speed71 tok/sMediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 24.4s; list price $4/20 per 1M in/out
Artificial Analysis output speed73 tok/sHighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 38.9s; list price $4/20 per 1M in/out
Artificial Analysis output speed77 tok/sExtra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 137.2s; list price $4/20 per 1M in/out
Artificial Analysis output speed90 tok/sMaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 724.5s; list price $4/20 per 1M in/out
ArXivMath (no tools)91.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—ArXivMath Aug 2026 (57 problems), no tools, avg of 4
ArXivMath (with tools)96.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—Aug 2026, code execution
AutomationBench24.2%Loweffort=lowIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.51; competitor number as plotted by OpenAI ("Opus 5.5 w/ fallbacks")
AutomationBench23.3%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps); run by Zapier without fallback models (safeguard interventions counted as failures); vendor-reported cost/task $0.5
AutomationBench29.5%Mediumeffort=mediumIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.65; competitor number as plotted by OpenAI ("Opus 5.5 w/ fallbacks")
AutomationBench28.6%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps); run by Zapier without fallback models (safeguard interventions counted as failures); vendor-reported cost/task $0.64
AutomationBench33.0%Higheffort=highIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.71; competitor number as plotted by OpenAI ("Opus 5.5 w/ fallbacks")
AutomationBench32.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps); run by Zapier without fallback models (safeguard interventions counted as failures); vendor-reported cost/task $0.7
AutomationBench35.8%Extra higheffort=xhighIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗29 Sep 2026—AutomationBench 1.0.6 (Zapier); vendor-estimated API cost/task $0.89; competitor number as plotted by OpenAI ("Opus 5.5 w/ fallbacks")
AutomationBench34.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps); run by Zapier without fallback models (safeguard interventions counted as failures); vendor-reported cost/task $0.86
AutomationBench40.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—AutomationBench (Zapier; end-to-end business workflows across connected apps); run by Zapier without fallback models (safeguard interventions counted as failures); vendor-reported cost/task $1.37
Chartography (no tools)64.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗28 Sep 2026—no tools
Chartography (with tools)89.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—Chartography (Surge AI chart-reading) with tools
CursorBench43.7%LowOpus 5.5 LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 18; $1.17/task; 15,811 tokens/task; 28 steps/task
CursorBench43.7%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; Opus 5.5 cost estimated by Anthropic from Cursor token counts; vendor-reported cost/task $1.18
CursorBench52.5%MediumOpus 5.5 MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 6; $2.91/task; 37,954 tokens/task; 54 steps/task
CursorBench52.5%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; Opus 5.5 cost estimated by Anthropic from Cursor token counts; vendor-reported cost/task $2.9
CursorBench56.0%HighOpus 5.5 HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 3; $3.97/task; 53,078 tokens/task; 68 steps/task
CursorBench56.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; Opus 5.5 cost estimated by Anthropic from Cursor token counts; vendor-reported cost/task $3.97
CursorBench56.0%Extra highOpus 5.5 Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 2; $6.98/task; 101,083 tokens/task; 109 steps/task
CursorBench56.0%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; Opus 5.5 cost estimated by Anthropic from Cursor token counts; vendor-reported cost/task $6.99
CursorBench57.8%MaxOpus 5.5 MaxIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 1; $13.43/task; 218,363 tokens/task; 185 steps/task
CursorBench57.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Cursor agentCursorBench 4.0, run end-to-end by Cursor in its production agent harness; Opus 5.5 cost estimated by Anthropic from Cursor token counts; vendor-reported cost/task $13.43
DeepSWE68.4%MaxOpus 5.5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE74.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—avg of 5 trials unless otherwise noted
Design Arena (all categories)1399Defaultclaude-opus-5-5Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±10.4 SE; 1259 battles; win rate 65.3%
Epoch Capabilities Index167.3DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 164.0-172.0; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)2050Default*claude-opus-5-5Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 7; rubric score 16.79/20; slop 10.43; avg length 6041 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*)
FrontierCode47.3%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Anthropic: scores decline above medium and mostly recover at max (out-of-scope changes penalized); vendor-reported cost/task $0.4
FrontierCode54.6%Mediumclaude-opus-5-5_mediumIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026Claude Coderead from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness claude-code; Mean@5
FrontierCode54.6%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Anthropic: scores decline above medium and mostly recover at max (out-of-scope changes penalized); vendor-reported cost/task $0.8
FrontierCode54.0%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Anthropic: scores decline above medium and mostly recover at max (out-of-scope changes penalized); vendor-reported cost/task $1.09
FrontierCode51.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Anthropic: scores decline above medium and mostly recover at max (out-of-scope changes penalized); vendor-reported cost/task $2.25
FrontierCode54.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Main (150 tasks, run & scored by Cognition; composite of correctness + mergeability rubric; penalizes out-of-scope edits); Anthropic: scores decline above medium and mostly recover at max (out-of-scope changes penalized); vendor-reported cost/task $6.19
FrontierCode v1.1 (Extended)65.3%Mediumadaptive thinking, effort=medium (best)Maker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude CodeFrontierCode v1.1 Extended; declines above medium, mostly recovers at max
FrontierCode v1.1 (Extended)63.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code
FrontierMath (Tiers 1-3)91.2%Maxclaude-opus-5-5_maxIndependent testIndependentEpoch AI ↗22 Sep 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 1.7pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath Tier 495.0%Maxclaude-opus-5-5_maxIndependent testIndependentEpoch AI ↗22 Sep 2026—FrontierMath Tier 4 (v2), Epoch-run; stderr 4.1pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tier_4_v2.csv)
FrontierSWE62.3%DefaultIndependent testIndependentFrontierSWE ↗1 Oct 2026proximusFrontierSWE V2, mean@5 over 34 tasks (20h budget); ±9.8; $98.87/trial; 14.2h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort
FrontierSWE v262.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Proximal harnessavg of 5 trials unless otherwise noted
GDP.pdf25.6%Loweffort=lowIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.76412; competitor number as plotted by OpenAI ("Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback")
GDP.pdf25.6%Mediumeffort=mediumIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.79532; competitor number as plotted by OpenAI ("Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback")
GDP.pdf28.8%Higheffort=highIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.82543; competitor number as plotted by OpenAI ("Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback")
GDP.pdf26.6%Extra higheffort=xhighIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $0.96427; competitor number as plotted by OpenAI ("Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback")
GDP.pdf26.2%Maxeffort=maxIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗29 Sep 2026—GDP.pdf (Surge AI; professional questions over complex PDFs); vendor-estimated API cost/task $1.55234; competitor number as plotted by OpenAI ("Claude Opus 5.5 w/ Claude Opus 5 / Claude Opus 4.8 fallback")
GDPval-AA v2.11224Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $0.21
GDPval-AA v2.11576Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $0.86
GDPval-AA v2.11692Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $1.54
GDPval-AA v2.11820Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1798.02-1842.19
GDPval-AA v2.11820Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $4.21
GDPval-AA v2.11846MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1822.93-1869.4
GDPval-AA v2.11846Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—GDPval-AA v2.1 (Artificial Analysis, Elo from blind pairwise comparisons, 44 occupations); vendor-reported cost/task $8.92
Harvey LAB-AA91.2%MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
HealthBench Professional65.6%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—avg of 5 trials unless otherwise noted
HLE Diamond55.0%Defaultclaude-opus-5-5Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±3.1; entry added 2026-09-03
Humanity's Last Exam48.3%LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-low)
Humanity's Last Exam54.7%MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-medium)
Humanity's Last Exam55.6%HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-high)
Humanity's Last Exam57.5%Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-xhigh)
Humanity's Last Exam61.4%MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5)
Humanity's Last Exam64.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—avg of 5 trials unless otherwise noted
Humanity's Last Exam (with tools)67.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—with web search, web fetch, programmatic tool calling, code execution; 1M total token cap
IOI (Vals)95.1%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.944 stderr; $5.258/test
LLM Creative Story-Writing Benchmark (Lech Mazur)3.8HighClaude Opus 5.5 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 3/56; Thurstone comparison score (centered at 0); est. win chance 93%; 95% bootstrap 3.676 to 3.882
LMArena Code Arena (WebDev)1818Maxclaude-opus-5.5-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 1 (CI rank 1-1); 95% CI 1801-1834; 1976 votes
LMArena Text - Coding category1538Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 11 (CI rank 1-50); 95% CI 1519-1557; 918 votes
LMArena Text - Creative Writing1515Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 2 (rank range 1-7); 95% CI 1493.9-1536.1; 908 votes; style-controlled
LMArena Text - Expert1539Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 4 (rank range 1-43); 95% CI 1511.1-1566.8; 448 votes; style-controlled
LMArena Text - Hard Prompts1533Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 3 (rank range 1-11); 95% CI 1521.1-1545.3; 2437 votes; style-controlled
LMArena Text - Instruction Following1514Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 2 (rank range 1-10); 95% CI 1498.2-1530.7; 1337 votes; style-controlled
LMArena Text - Longer Query1528Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 2 (rank range 1-9); 95% CI 1512.9-1542.7; 1708 votes; style-controlled
LMArena Text - Multi-Turn1497Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 11 (rank range 2-75); 95% CI 1470.6-1523.1; 496 votes; style-controlled
LMArena Text - Non-English1496Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 2 (rank range 1-16); 95% CI 1484.2-1508.1; 2470 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1493Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 10 (rank range 1-55); 95% CI 1470.7-1515.0; 716 votes; style-controlled
LMArena Text - Occupational: Legal & Government1490Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 23 (rank range 1-100); 95% CI 1457.7-1523.0; 336 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1519Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 2 (rank range 1-7); 95% CI 1500.8-1537.3; 1159 votes; style-controlled
LMArena Text (overall)1504Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 4 (CI rank 2-14); 95% CI 1494-1514; 3932 votes
OfficeQA Pro67.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—avg of 5 trials unless otherwise noted
OSWorld 2.1 (partial score)81.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—partial score; system card table labels it OSWorld 2.0 (partial/strict 81.8/48.7)
OSWorld 2.1 (strict pass rate)48.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—strict pass rate (card labels OSWorld 2.0)
PRBench Legal (Scale)47.6%Defaultclaude-opus-5-5Independent testIndependentScale AI (SEAL) ↗25 Sep 2026—Scale rank 8; ±2.2 CI
ProgramBench (fully resolved)18.5%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±2.753 stderr; $68.051/test; strict fully-resolved rate
ProgramBench (fully resolved)91.2%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026mini-swe-agentavg of 5 trials unless otherwise noted
SciCode58.6%LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-low)
SciCode59.3%MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-medium)
SciCode60.4%HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-high)
SciCode65.0%Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-xhigh)
SciCode66.9%MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5)
SimpleBench88.4%DefaultClaude Opus 5.5Independent testIndependentSimpleBench ↗24 Sep 2026—AVG@5, temp 0.7; rank 1st; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified72.2%Maxclaude-opus-5-5_maxIndependent testIndependentEpoch AI ↗22 Sep 2026—Epoch-run (no tools); ±1.42 stderr
SWE Atlas - Codebase QnA66.4%MaxOpus 5.5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE-bench Multilingual93.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—avg of 5 trials unless otherwise noted
SWE-bench Multimodal61.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—avg of 5 trials unless otherwise noted
SWE-bench Pro (Anthropic internal subset)87.4%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗1 Oct 2026—Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $0.12 per solved task
SWE-bench Pro (Anthropic internal subset)92.8%Mediumadaptive thinking, effort=medium (default)Maker's own figureVendor-reportedAnthropic ↗1 Oct 2026—Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $0.22 per solved task; ~2.5 pts below high for ~70% of the cost
SWE-bench Pro (Anthropic internal subset)95.3%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗1 Oct 2026—Anthropic internal 478-problem SWE-bench Pro subset (not comparable to public leaderboard), priced as billed; $0.29 per task; xhigh adds ~1.4 pts for 2.5x the cost of high
SWE-Bench Pro (public, v1)89.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—avg of 5 trials unless otherwise noted
Terminal-Bench 2.187.6%Higheffort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±1.716 stderr; $0.496/test
Terminal-Bench 4.031.3%LowClaude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-low)
Terminal-Bench 4.038.5%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; xhigh is the reported headline (max within noise, SE ±2.6); production safeguards on, fallbacks on 2.5% of requests; vendor-reported cost/task $1.29
Terminal-Bench 4.052.5%MediumClaude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-medium)
Terminal-Bench 4.057.6%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; xhigh is the reported headline (max within noise, SE ±2.6); production safeguards on, fallbacks on 2.5% of requests; vendor-reported cost/task $2.94
Terminal-Bench 4.056.6%HighClaude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-high)
Terminal-Bench 4.064.2%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; xhigh is the reported headline (max within noise, SE ±2.6); production safeguards on, fallbacks on 2.5% of requests; vendor-reported cost/task $3.88
Terminal-Bench 4.059.6%Extra highClaude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5-xhigh)
Terminal-Bench 4.066.4%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; xhigh is the reported headline (max within noise, SE ±2.6); production safeguards on, fallbacks on 2.5% of requests; vendor-reported cost/task $7.35
Terminal-Bench 4.065.2%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0 stderr; $13.199/test
Terminal-Bench 4.063.1%MaxOpus 5.5 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.059.6%MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug claude-opus-5-5)
Terminal-Bench 4.064.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench 4.0 (66 tasks); Anthropic internal runs, 5-15 trials/task; xhigh is the reported headline (max within noise, SE ±2.6); production safeguards on, fallbacks on 2.5% of requests; vendor-reported cost/task $11.24
Terminal-Bench-Science 0.163.3%Maxeffort=maxIndependent testIndependentOpenAI (competitor result in OpenAI launch post) ↗29 Sep 2026—Terminal-Bench Science 0.1; vendor-estimated API cost/task $23.2108; competitor number as plotted by OpenAI
Terminal-Bench-Science 0.158.7%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026Claude Code (--bare)Terminal-Bench-Science 0.1 (70 tasks); SE ±4.8; safeguards on (fallbacks on 3.9% of requests)
Toolathlon77.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—avg of 5 trials unless otherwise noted
Vals Code Migration66.7%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.329 stderr; $112.971/test
Vals Index67.0%Maxeffort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±0.888 stderr; $32.144/test
Vals Legal Research Bench50.5%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-opus-5-5; rank 5/72; ±3.475 stderr; $25.163725/test
Vals Public Benefits Bench v1.170.6%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id anthropic/claude-opus-5-5; rank 3/45; ±1.185 stderr; $5.952407/test
Vals SRE Bench33.6%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±2.923 stderr; $34.654/test
Vals Vibe Code Bench90.3%Maxeffort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±1.526 stderr; $57.923/test
Vending-Bench 2$9,235DefaultIndependent testIndependentAndon Labs ↗1 Oct 2026—final money balance after simulated year, arithmetic mean across runs; ±$785; rank 8; only top 10 rendered server-side (57 more behind 'Show more')
WANDR31.2%Lowadaptive thinking, effort=lowMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $1.2
WANDR62.8%Mediumadaptive thinking, effort=medMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $11.2
WANDR67.3%Highadaptive thinking, effort=highMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $16.88
WANDR71.3%Extra highadaptive thinking, effort=xhighMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $29.06
WANDR72.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗22 Sep 2026—WANDR (Perplexity wide-search benchmark), soft F1; Anthropic-modified setup (offline search/fetch, 980k token budget), not comparable to Perplexity numbers; vendor-reported cost/task $37.92