Skip to content
Bencher

TestsBenchmarks · Coding

SWE-Bench Pro V2 (full)

Refreshed 642-task SWE-Bench Pro public split with a locked, network-isolated protocol and re-grading on a pristine image.

% solvedHigher is betterToo easy now: the top models all score near perfectSaturatedThe test's websiteOfficial page ↗

The best 10 · 10 models tested

Top 10 · best setting per model · 10 models, 10 results (10 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
SWE-Bench Pro V2 (full) leaderboard
Thinking levelSetting
1Claude Opus 5Anthropic99.4%Extra highOpus 5 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗Claude Code1 Oct 2026
2Claude Fable 5.1Anthropic99.1%HighFable 5.1 (Claude Code) highIndependent testIndependentScale AI SEAL ↗Claude Code1 Oct 2026
3Kimi K3Moonshot AI (Kimi)97.7%MaxKimi-K3 (mini-swe-agent) maxIndependent testIndependentScale AI SEAL ↗mini-SWE-agent1 Oct 2026
4GPT-6 AstraOpenAI96.9%HighGPT-6-Astra (Codex) highIndependent testIndependentScale AI SEAL ↗Codex1 Oct 2026
5GLM-5.3Z.ai (Zhipu)95.6%MaxGLM-5.3 (mini-swe-agent) maxIndependent testIndependentScale AI SEAL ↗mini-SWE-agent1 Oct 2026
6GPT-5.6 SolOpenAI95.5%Extra highGPT-5.6-sol (Codex) xhighIndependent testIndependentScale AI SEAL ↗Codex1 Oct 2026
7Gemini 3.8 FlashGoogle (Gemini / DeepMind)94.9%HighGemini 3.8 Flash (mini-swe-agent) highIndependent testIndependentScale AI SEAL ↗mini-SWE-agent1 Oct 2026
8Claude Sonnet 5Anthropic93.2%Extra highSonnet 5 (Claude Code) xhighIndependent testIndependentScale AI SEAL ↗Claude Code1 Oct 2026
9GPT-5.6 TerraOpenAI92.4%Extra highGPT-5.6-Terra (Codex) xhighIndependent testIndependentScale AI SEAL ↗Codex1 Oct 2026
10InklingThinking Machines Lab89.9%Extra highInkling (mini-swe-agent) xhighIndependent testIndependentScale AI SEAL ↗mini-SWE-agent1 Oct 2026