TestsBenchmarks · Coding
SWE-Bench Pro V2 (full)
Refreshed 642-task SWE-Bench Pro public split with a locked, network-isolated protocol and re-grading on a pristine image.
% solvedHigher is betterToo easy now: the top models all score near perfectSaturatedThe test's websiteOfficial page ↗
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 99.4% | Extra highOpus 5 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | Claude Code | 1 Oct 2026 |
| 2 | Claude Fable 5.1 | Anthropic | 99.1% | HighFable 5.1 (Claude Code) high | Independent testIndependentScale AI SEAL ↗ | Claude Code | 1 Oct 2026 |
| 3 | Kimi K3 | Moonshot AI (Kimi) | 97.7% | MaxKimi-K3 (mini-swe-agent) max | Independent testIndependentScale AI SEAL ↗ | mini-SWE-agent | 1 Oct 2026 |
| 4 | GPT-6 Astra | OpenAI | 96.9% | HighGPT-6-Astra (Codex) high | Independent testIndependentScale AI SEAL ↗ | Codex | 1 Oct 2026 |
| 5 | GLM-5.3 | Z.ai (Zhipu) | 95.6% | MaxGLM-5.3 (mini-swe-agent) max | Independent testIndependentScale AI SEAL ↗ | mini-SWE-agent | 1 Oct 2026 |
| 6 | GPT-5.6 Sol | OpenAI | 95.5% | Extra highGPT-5.6-sol (Codex) xhigh | Independent testIndependentScale AI SEAL ↗ | Codex | 1 Oct 2026 |
| 7 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 94.9% | HighGemini 3.8 Flash (mini-swe-agent) high | Independent testIndependentScale AI SEAL ↗ | mini-SWE-agent | 1 Oct 2026 |
| 8 | Claude Sonnet 5 | Anthropic | 93.2% | Extra highSonnet 5 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | Claude Code | 1 Oct 2026 |
| 9 | GPT-5.6 Terra | OpenAI | 92.4% | Extra highGPT-5.6-Terra (Codex) xhigh | Independent testIndependentScale AI SEAL ↗ | Codex | 1 Oct 2026 |
| 10 | Inkling | Thinking Machines Lab | 89.9% | Extra highInkling (mini-swe-agent) xhigh | Independent testIndependentScale AI SEAL ↗ | mini-SWE-agent | 1 Oct 2026 |