TestsBenchmarks · Coding
SWE-Bench Pro V2 (hard)
Hard subset of SWE-Bench Pro V2; still separates frontier agents better than the full set.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 98.0% | Extra highOpus 5 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | Claude Code | 1 Oct 2026 |
| 2 | Claude Fable 5.1 | Anthropic | 92.2% | HighFable 5.1 (Claude Code) high | Independent testIndependentScale AI SEAL ↗ | Claude Code | 1 Oct 2026 |
| 3 | GPT-6 Astra | OpenAI | 90.2% | HighGPT-6 Astra (Codex) high | Independent testIndependentScale AI SEAL ↗ | Codex | 1 Oct 2026 |
| 4 | Claude Sonnet 5 | Anthropic | 88.2% | Extra highSonnet 5 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | Claude Code | 1 Oct 2026 |
| 5 | Kimi K3 | Moonshot AI (Kimi) | 88.2% | MaxKimi-K3 (mini-swe-agent) max | Independent testIndependentScale AI SEAL ↗ | mini-SWE-agent | 1 Oct 2026 |
| 6 | GPT-5.6 Terra | OpenAI | 86.3% | Extra highGPT-5.6 Terra (Codex) xhigh | Independent testIndependentScale AI SEAL ↗ | Codex | 1 Oct 2026 |
| 7 | GLM-5.3 | Z.ai (Zhipu) | 84.3% | MaxGLM-5.3 (mini-swe-agent) max | Independent testIndependentScale AI SEAL ↗ | mini-SWE-agent | 1 Oct 2026 |
| 8 | GPT-5.6 Sol | OpenAI | 82.4% | Extra highGPT-5.6 Sol (Codex) xhigh | Independent testIndependentScale AI SEAL ↗ | Codex | 1 Oct 2026 |
| 9 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 58.8% | HighGemini 3.8 Flash (mini-swe-agent) high | Independent testIndependentScale AI SEAL ↗ | mini-SWE-agent | 1 Oct 2026 |
| 10 | Inkling | Thinking Machines Lab | 56.9% | Extra highInkling (mini-swe-agent) xhigh | Independent testIndependentScale AI SEAL ↗ | mini-SWE-agent | 1 Oct 2026 |
| 11 | Claude Haiku 4.5 | Anthropic | 25.5% | Extra highHaiku 4.5 (Claude Code) xhigh | Independent testIndependentScale AI SEAL ↗ | Claude Code | 1 Oct 2026 |