TestsBenchmarks · Coding
SWE-rebench
Resolve rate on fresh, decontaminated GitHub issues collected in a rolling time window, run with one standard scaffold.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 64.5% | HighFable 5 [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 2 | Grok 4.5 | SpaceXAI (formerly xAI) | 63.8% | HighGrok 4.5 [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 3 | Claude Opus 5 | Anthropic | 63.4% | HighOpus 5 [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 4 | GLM-5.2 | Z.ai (Zhipu) | 62.9% | HighGLM-5.2 [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 5 | GPT-5.6 Sol | OpenAI | 62.3% | MediumGPT-5.6 Sol [medium] | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 6 | Claude Sonnet 5 | Anthropic | 56.8% | HighSonnet 5 [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 7 | MiniMax M3 | MiniMax | 47.2% | DefaultMiniMax M3 | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 8 | MiMo-V2.5-Pro | Xiaomi MiMo | 46.5% | DefaultMiMo V2.5 Pro | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 9 | GPT-5.6 Luna | OpenAI | 43.6% | MediumGPT-5.6 Luna [medium] | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 10 | DeepSeek V4 Pro (Preview, 0423) | DeepSeek | 40.2% | HighDeepSeek-V4 Pro [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 11 | Qwen3.6 27B | Qwen (Alibaba) | 31.2% | DefaultQwen3.6-27B | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 12 | Qwen3.6 35B A3B | Qwen (Alibaba) | 24.7% | DefaultQwen3.6-35B-A3B | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |
| 13 | Qwen3.5-35B-A3B | Qwen (Alibaba) | 17.1% | DefaultQwen3.5-35B-A3B | Independent testIndependentSWE-rebench (Nebius) ↗ | SWE-rebench standard scaffold | 1 Oct 2026 |