TestsBenchmarks · Coding
SWE-bench Verified
Share of 500 human-validated real GitHub Python issues a model fixes so the repo tests pass.
% solvedHigher is betterToo easy now: the top models all score near perfectSaturatedThe test's websiteOfficial page ↗
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 97.0% | Default | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 2 | DeepSeek V4 Pro (0813) | DeepSeek | 96.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 3 | GPT-5.6 Sol | OpenAI | 96.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 4 | Grok 4.6 | SpaceXAI (formerly xAI) | 95.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 5 | GPT-5.6 Terra | OpenAI | 95.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 6 | GLM-5.3 | Z.ai (Zhipu) | 95.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 7 | Claude Fable 5 | Anthropic | 95.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 8 | Kimi K3 | Moonshot AI (Kimi) | 93.4% | Default | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 9 | GPT-5.6 Luna | OpenAI | 93.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 10 | GLM-5.3-Flash | Z.ai (Zhipu) | 92.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 11 | DeepSeek V4 Flash (0731) | DeepSeek | 88.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 12 | Claude Opus 4.8 | Anthropic | 88.6% | Maxeffort=max | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 13 | Grok 4.5 | SpaceXAI (formerly xAI) | 86.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 14 | Muse Spark 1.2 | Meta (Meta Superintelligence Labs, Muse) | 86.6% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 15 | Qwen3.8 27B | Qwen (Alibaba) | 86.0% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 16 | Qwen3.8 Max (0902) | Qwen (Alibaba) | 85.6% | Default | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 17 | Claude Sonnet 5 | Anthropic | 85.2% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 30 Jun 2026 |
| 18 | GLM-5.2 | Z.ai (Zhipu) | 82.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 19 | GPT-5.5 | OpenAI | 82.6% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 20 | Inkling Small | Thinking Machines Lab | 82.2% | Defaultreasoning_effort=0.99 | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 21 | Claude Opus 4.7 | Anthropic | 82.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 22 | Muse Spark 1.1 | Meta (Meta Superintelligence Labs, Muse) | 82.0% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 23 | Gemini 3.7 Flash | Google (Gemini / DeepMind) | 80.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 24 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 80.6% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | — | 19 Feb 2026 |
| 25 | MiniMax M3 | MiniMax | 80.5% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | Claude Code | 1 Jun 2026 |
| 26 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 80.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | bash-only (single bash tool) agent | 1 Sep 2026 |
| 27 | MiniMax M2.7 | MiniMax | 79.9% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | Claude Code | 1 Jun 2026 |
| 28 | Composer 2.5 | Cursor | 79.6% | Default | Independent testIndependentVals.ai ↗ | Cursor CLI | 1 Sep 2026 |
| 29 | DeepSeek V4 Pro (Preview, 0423) | DeepSeek | 79.4% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | — | 24 Apr 2026 |
| 30 | Gemini 3.5 Flash | Google (Gemini / DeepMind) | 79.3% | Highgemini-3.5-flash_high | Independent testIndependentEpoch AI ↗ | — | 1 Jun 2026 |
| 31 | DeepSeek V4 Flash (Preview, 0423) | DeepSeek | 79.0% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | — | 24 Apr 2026 |
| 32 | Qwen3.6 Plus | Qwen (Alibaba) | 78.8% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | Internal scaffold (bash + file-edit) | 2 Apr 2026 |
| 33 | Claude Opus 4.6 | Anthropic | 78.7% | Defaultclaude-opus-4-6 | Independent testIndependentEpoch AI ↗ | — | 18 Feb 2026 |
| 34 | Qwen3.7 Plus | Qwen (Alibaba) | 77.7% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | Internal scaffold (bash + file-edit) | 21 May 2026 |
| 35 | Mistral Medium 3.5 | Mistral AI | 77.6% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | — | 22 May 2026 |
| 36 | Qwen3.7 Max | Qwen (Alibaba) | 77.3% | Defaultqwen3.7-max | Independent testIndependentEpoch AI ↗ | — | 18 Jun 2026 |
| 37 | Qwen3.6 27B | Qwen (Alibaba) | 77.2% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | Internal scaffold (bash + file-edit) | 22 Apr 2026 |
| 38 | GPT-5.4 | OpenAI | 76.9% | Highgpt-5.4-2026-03-05_high | Independent testIndependentEpoch AI ↗ | — | 6 Mar 2026 |
| 39 | Kimi K2.6 | Moonshot AI (Kimi) | 76.7% | Defaultkimi-k2.6 | Independent testIndependentEpoch AI ↗ | — | 8 May 2026 |
| 40 | Qwen3.6 Max Preview | Qwen (Alibaba) | 76.7% | Defaultqwen3.6-max-preview | Independent testIndependentEpoch AI ↗ | — | 28 May 2026 |
| 41 | Muse Glimmer 30B | Meta (Meta Superintelligence Labs, Muse) | 76.0% | HighHigh reasoning | Maker's own figureVendor-reportedMeta ↗ | — | 10 Aug 2026 |
| 42 | Gemini 3.1 Pro Preview Custom Tools | Google (Gemini / DeepMind) | 75.6% | Defaultgemini-3.1-pro-preview-customtools | Independent testIndependentEpoch AI ↗ | — | 24 Feb 2026 |
| 43 | Claude Sonnet 4.6 | Anthropic | 75.2% | Defaultclaude-sonnet-4-6 | Independent testIndependentEpoch AI ↗ | — | 21 Feb 2026 |
| 44 | GPT-5.3-Codex | OpenAI | 74.8% | Highgpt-5.3-codex_high | Independent testIndependentEpoch AI ↗ | — | 25 Feb 2026 |
| 45 | GLM-5.1 | Z.ai (Zhipu) | 74.2% | Defaultglm-5.1 | Independent testIndependentEpoch AI ↗ | — | 15 May 2026 |
| 46 | Kimi K2.5 | Moonshot AI (Kimi) | 73.8% | Defaultkimi-k2.5 | Independent testIndependentEpoch AI ↗ | — | 17 Feb 2026 |
| 47 | Devstral 2 | Mistral AI | 72.2% | Default | Maker's own figureVendor-reportedMistral AI ↗ | — | 22 May 2026 |
| 48 | GLM 5 | Z.ai (Zhipu) | 72.1% | Defaultglm-5 | Independent testIndependentEpoch AI ↗ | — | 15 Feb 2026 |