Skip to content
Bencher

TestsBenchmarks · Coding

SWE-Bench Pro V2 (hard)

Hard subset of SWE-Bench Pro V2; still separates frontier agents better than the full set.

% solvedHigher is betterCounts toward:Weights: Coding ×0.8The test's websiteOfficial page ↗

The best 11 · 11 models tested

Top 11 · best setting per model · 11 models, 11 results (11 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
SWE-Bench Pro V2 (hard) leaderboard
Thinking levelSetting
1Claude Opus 5Anthropic98.0%Extra highOpus 5 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗Claude Code1 Oct 2026
2Claude Fable 5.1Anthropic92.2%HighFable 5.1 (Claude Code) highIndependent testIndependentScale AI SEAL ↗Claude Code1 Oct 2026
3GPT-6 AstraOpenAI90.2%HighGPT-6 Astra (Codex) highIndependent testIndependentScale AI SEAL ↗Codex1 Oct 2026
4Claude Sonnet 5Anthropic88.2%Extra highSonnet 5 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗Claude Code1 Oct 2026
5Kimi K3Moonshot AI (Kimi)88.2%MaxKimi-K3 (mini-swe-agent) maxIndependent testIndependentScale AI SEAL ↗mini-SWE-agent1 Oct 2026
6GPT-5.6 TerraOpenAI86.3%Extra highGPT-5.6 Terra (Codex) xhighIndependent testIndependentScale AI SEAL ↗Codex1 Oct 2026
7GLM-5.3Z.ai (Zhipu)84.3%MaxGLM-5.3 (mini-swe-agent) maxIndependent testIndependentScale AI SEAL ↗mini-SWE-agent1 Oct 2026
8GPT-5.6 SolOpenAI82.4%Extra highGPT-5.6 Sol (Codex) xhighIndependent testIndependentScale AI SEAL ↗Codex1 Oct 2026
9Gemini 3.8 FlashGoogle (Gemini / DeepMind)58.8%HighGemini 3.8 Flash (mini-swe-agent) highIndependent testIndependentScale AI SEAL ↗mini-SWE-agent1 Oct 2026
10InklingThinking Machines Lab56.9%Extra highInkling (mini-swe-agent) xhighIndependent testIndependentScale AI SEAL ↗mini-SWE-agent1 Oct 2026
11Claude Haiku 4.5Anthropic25.5%Extra highHaiku 4.5 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗Claude Code1 Oct 2026