TestsBenchmarks · Coding
ProgramBench (fully resolved)
Share of 200 programs (from jq to SQLite/FFmpeg) an agent re-implements from only the binary and docs so that all hidden behavioral tests pass.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5.5 | Anthropic | 79.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 28 Sep 2026 |
| 2 | Kimi K3 | Moonshot AI (Kimi) | 77.8% | Maxreasoning_effort=max | Maker's own figureVendor-reportedMoonshot AI ↗ | Kimi Code | 16 Jul 2026 |
| 3 | Kimi K2.7 Code | Moonshot AI (Kimi) | 53.6% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | Kimi Code CLI | 12 Jun 2026 |
| 4 | Kimi K2.6 | Moonshot AI (Kimi) | 48.3% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | Kimi Code CLI | 12 Jun 2026 |
| 5 | Claude Opus 5.5 | Anthropic | 18.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 6 | Claude Fable 5.1 | Anthropic | 7.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 7 | GPT-6 Astra | OpenAI | 5.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 8 | Claude Opus 5 | Anthropic | 4.5% | Extra high | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 9 | Gemini 4 Argon | Google (Gemini / DeepMind) | 2.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 10 | Muse Spark 1.3 | Meta (Meta Superintelligence Labs, Muse) | 2.5% | Max | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 11 | Claude Fable 5 | Anthropic | 2.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 12 | GPT-6 Sol | OpenAI | 2.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 13 | GPT-5.6 Sol | OpenAI | 1.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 14 | GLM-5.3 | Z.ai (Zhipu) | 1.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 15 | Claude Opus 4.8 | Anthropic | 1.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 16 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 1.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 17 | GLM-5.2 | Z.ai (Zhipu) | 0.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 18 | Grok 4.7 | SpaceXAI (formerly xAI) | 0.5% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 19 | MiMo-V2.6-Flash | Xiaomi MiMo | 0.5% | Default | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 20 | GPT-5.6 Terra | OpenAI | 0.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 21 | MiMo-V2.6-Pro | Xiaomi MiMo | 0.5% | Default | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 22 | GPT-5.5 | OpenAI | 0.5% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 23 | GPT-6 Luna | OpenAI | 0.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 24 | GPT-5.4 | OpenAI | 0.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 25 | Claude Sonnet 4.6 | Anthropic | 0.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 26 | Inkling Small | Thinking Machines Lab | 0.5% | Defaultreasoning_effort=0.99 | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 27 | Gemini 3.6 Flash | Google (Gemini / DeepMind) | 0.5% | Default | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 28 | Claude Sonnet 5 | Anthropic | 0.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 29 | Hy4 preview | Tencent Hunyuan | 0.0% | Default | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 30 | DeepSeek V4 Pro (0813) | DeepSeek | 0.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | mini-SWE-agent | 29 Sep 2026 |
| 31 | Gemini 3.7 Flash | Google (Gemini / DeepMind) | 0.0% | Default | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 32 | Muse Spark 1.2 | Meta (Meta Superintelligence Labs, Muse) | 0.0% | Extra high | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 33 | Claude Opus 4.7 | Anthropic | 0.0% | Extra high | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 34 | Gemini 3.5 Flash | Google (Gemini / DeepMind) | 0.0% | Default | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 35 | Claude Opus 4.6 | Anthropic | 0.0% | Default | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 36 | Muse Spark 1.1 | Meta (Meta Superintelligence Labs, Muse) | 0.0% | Extra high | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 37 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 0.0% | Default | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 38 | Gemini 3 Flash Preview | Google (Gemini / DeepMind) | 0.0% | Default | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 39 | Claude Haiku 4.5 | Anthropic | 0.0% | Default | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 40 | GPT-5.4 Mini | OpenAI | 0.0% | Default | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |
| 41 | GPT-5 Mini | OpenAI | 0.0% | Default | Independent testIndependentProgramBench ↗ | mini-SWE-agent | 28 Sep 2026 |