Skip to content
Bencher

TestsBenchmarks · Coding

FrontierSWE v2

Proximal's 34 ultra-long-horizon engineering/research tasks (~20h per task), run in Proximal's harness at max effort.

% solvedHigher is betterCounts toward:Weights: Coding ×0.6The test's websiteOfficial page ↗

The best 8 · 8 models tested

Top 8 · best setting per model · 8 models, 8 results (2 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
FrontierSWE v2 leaderboard
Thinking levelSetting
1GPT-6 AstraOpenAI65.5%Maxreasoning effort=maxIndependent testIndependentProximal (via Anthropic system card) ↗Proximal harness22 Sep 2026
2Claude Opus 5.5Anthropic62.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗Proximal harness22 Sep 2026
3Claude Sonnet 5.5Anthropic61.9%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗Proximal harness28 Sep 2026
4Claude Fable 5.1Anthropic56.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗Proximal harness22 Sep 2026
5Gemini 4 ArgonGoogle (Gemini / DeepMind)55.0%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗—30 Sep 2026
6Claude Opus 5Anthropic52.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗Proximal harness1 Sep 2026
7Claude Fable 5Anthropic48.0%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗Proximal harness1 Sep 2026
8GPT-5.6 SolOpenAI32.2%Maxreasoning effort=maxIndependent testIndependentProximal (via Anthropic system card) ↗Proximal harness22 Sep 2026