TestsBenchmarks · Reasoning
SimpleBench
Accuracy on trick multiple-choice questions about everyday spatial, temporal and social reasoning where ordinary humans do well.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | Anthropic | 88.4% | DefaultClaude Opus 5.5 | Independent testIndependentSimpleBench ↗ | — | 24 Sep 2026 |
| 2 | Claude Fable 5.1 | Anthropic | 86.6% | DefaultClaude Fable 5.1 | Independent testIndependentSimpleBench ↗ | — | 3 Sep 2026 |
| 3 | GPT-6 Astra Pro | OpenAI | 86.5% | DefaultGPT-6 Astra Pro | Independent testIndependentSimpleBench ↗ | — | 7 Sep 2026 |
| 4 | GPT-6 Astra | OpenAI | 83.6% | DefaultGPT-6 Astra | Independent testIndependentSimpleBench ↗ | — | 7 Sep 2026 |
| 5 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 82.4% | DefaultGemini 3.8 Flash | Independent testIndependentSimpleBench ↗ | — | 3 Sep 2026 |
| 6 | Claude Fable 5 | Anthropic | 81.9% | DefaultClaude Fable | Independent testIndependentSimpleBench ↗ | — | 10 Jun 2026 |
| 7 | Muse Spark 1.3 | Meta (Meta Superintelligence Labs, Muse) | 81.8% | DefaultMuse Spark 1.3 | Independent testIndependentSimpleBench ↗ | — | 3 Sep 2026 |
| 8 | Claude Opus 5 | Anthropic | 80.6% | DefaultClaude Opus 5 | Independent testIndependentSimpleBench ↗ | — | 24 Jul 2026 |
| 9 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 79.6% | DefaultGemini 3.1 Pro Preview | Independent testIndependentSimpleBench ↗ | — | 17 Feb 2026 |
| 10 | GPT-5.5 Pro | OpenAI | 76.9% | DefaultGPT-5.5 Pro | Independent testIndependentSimpleBench ↗ | — | 24 Apr 2026 |
| 11 | Gemini 3.5 Flash | Google (Gemini / DeepMind) | 76.7% | DefaultGemini 3.5 Flash | Independent testIndependentSimpleBench ↗ | — | 20 May 2026 |
| 12 | Grok 4.6 | SpaceXAI (formerly xAI) | 75.9% | DefaultGrok 4.6 | Independent testIndependentSimpleBench ↗ | — | 13 Aug 2026 |
| 13 | Muse Spark 1.2 | Meta (Meta Superintelligence Labs, Muse) | 74.5% | DefaultMuse Spark 1.2 | Independent testIndependentSimpleBench ↗ | — | 13 Aug 2026 |
| 14 | GPT-6 Sol | OpenAI | 73.1% | DefaultGPT-6 Sol | Independent testIndependentSimpleBench ↗ | — | 24 Sep 2026 |
| 15 | GPT-5.6 Sol Pro | OpenAI | 71.7% | Extra highGPT-5.6 Sol Pro (xhigh) | Independent testIndependentSimpleBench ↗ | — | 9 Jul 2026 |
| 16 | Qwen3.7 Max | Qwen (Alibaba) | 70.4% | DefaultQwen 3.7 Max | Independent testIndependentSimpleBench ↗ | — | 22 May 2026 |
| 17 | Grok 4.5 | SpaceXAI (formerly xAI) | 70.0% | DefaultGrok 4.5 | Independent testIndependentSimpleBench ↗ | — | 8 Jul 2026 |
| 18 | GPT-5.5 | OpenAI | 69.0% | DefaultGPT-5.5 | Independent testIndependentSimpleBench ↗ | — | 24 Apr 2026 |
| 19 | Claude Opus 4.6 | Anthropic | 67.6% | DefaultClaude Opus 4.6 | Independent testIndependentSimpleBench ↗ | — | 17 Feb 2026 |
| 20 | DeepSeek V4.1 Flash | DeepSeek | 66.7% | DefaultDeepSeek V4.1 Flash | Independent testIndependentSimpleBench ↗ | — | 12 Sep 2026 |
| 21 | GLM-5.3 | Z.ai (Zhipu) | 66.2% | DefaultGLM 5.3 | Independent testIndependentSimpleBench ↗ | — | 19 Aug 2026 |
| 22 | Claude Opus 4.8 | Anthropic | 64.8% | DefaultClaude Opus 4.8 | Independent testIndependentSimpleBench ↗ | — | 29 May 2026 |
| 23 | GPT-5.6 Sol | OpenAI | 64.8% | Extra highGPT-5.6 Sol (xhigh) | Independent testIndependentSimpleBench ↗ | — | 9 Jul 2026 |
| 24 | Qwen3.6 Max Preview | Qwen (Alibaba) | 63.0% | DefaultQwen 3.6 Max Preview | Independent testIndependentSimpleBench ↗ | — | 20 May 2026 |
| 25 | Qwen3.8 2.4T A95B (open weights) | Qwen (Alibaba) | 62.5% | DefaultQwen 3.8 2.4T A95B | Independent testIndependentSimpleBench ↗ | — | 13 Aug 2026 |
| 26 | Claude Opus 4.7 | Anthropic | 61.7% | DefaultClaude Opus 4.7 | Independent testIndependentSimpleBench ↗ | — | 22 Apr 2026 |
| 27 | DeepSeek V4 Flash (Preview, 0423) | DeepSeek | 61.1% | DefaultDeepSeek V4 Flash | Independent testIndependentSimpleBench ↗ | — | 3 Aug 2026 |
| 28 | Kimi K3 | Moonshot AI (Kimi) | 60.7% | MaxKimi K3 (max) | Independent testIndependentSimpleBench ↗ | — | 17 Jul 2026 |
| 29 | Claude Sonnet 5 | Anthropic | 60.6% | DefaultClaude Sonnet 5 | Independent testIndependentSimpleBench ↗ | — | 9 Jul 2026 |
| 30 | Qwen3.8 27B | Qwen (Alibaba) | 60.2% | DefaultQwen 3.8 27B | Independent testIndependentSimpleBench ↗ | — | 20 Aug 2026 |