TestsBenchmarks · Coding
LiveCodeBench
Pass@1 on recent competitive-programming problems from LeetCode, AtCoder and Codeforces (Vals implementation).
% solvedHigher is betterToo easy now: the top models all score near perfectSaturatedThe test's websiteOfficial page ↗
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8 Flash | Qwen (Alibaba) | 91.9% | Defaultthinking (setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | — | 26 Aug 2026 |
| 2 | DeepSeek V4 Flash (Preview, 0423) | DeepSeek | 91.6% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | — | 24 Apr 2026 |
| 3 | Claude Fable 5.1 | Anthropic | 90.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 4 | Qwen3.8 27B | Qwen (Alibaba) | 90.3% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | — | 14 Aug 2026 |
| 5 | DeepSeek V4 Pro (Preview, 0423) | DeepSeek | 89.8% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | — | 24 Apr 2026 |
| 6 | Claude Fable 5 | Anthropic | 89.8% | Maxeffort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 7 | Qwen3.7 Plus | Qwen (Alibaba) | 89.6% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | — | 21 May 2026 |
| 8 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 89.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 9 | Claude Opus 5 | Anthropic | 89.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 10 | Gemini 3.7 Flash | Google (Gemini / DeepMind) | 88.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 11 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 88.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 12 | Grok 4.6 | SpaceXAI (formerly xAI) | 88.2% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 13 | Gemini 3.6 Flash | Google (Gemini / DeepMind) | 88.1% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 14 | GPT-5.2-Codex | OpenAI | 88.0% | Default | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 15 | Qwen3.8 Max (0902) | Qwen (Alibaba) | 87.8% | Default | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 16 | Claude Opus 4.8 | Anthropic | 87.8% | Maxeffort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 17 | Gemini 3.5 Flash | Google (Gemini / DeepMind) | 87.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 18 | DeepSeek V4 Pro (0813) | DeepSeek | 87.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 19 | Grok 4.5 | SpaceXAI (formerly xAI) | 87.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 20 | GPT-5.3-Codex | OpenAI | 87.3% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 21 | DeepSeek V4 Flash (0731) | DeepSeek | 87.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 22 | Kimi K3 | Moonshot AI (Kimi) | 87.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 23 | Qwen3.7 Max | Qwen (Alibaba) | 87.1% | Default | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 24 | Kimi K2.6 | Moonshot AI (Kimi) | 86.8% | Default | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 25 | GPT-5 Mini | OpenAI | 86.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 26 | GPT-5.1 | OpenAI | 86.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 27 | Gemini 3 Pro | Google (Gemini / DeepMind) | 86.4% | Default | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 28 | Nemotron 3 Ultra | NVIDIA | 86.0% | Default | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 29 | Qwen3.6 Plus | Qwen (Alibaba) | 86.0% | Default | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 30 | Qwen3.6 27B | Qwen (Alibaba) | 83.9% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | — | 22 Apr 2026 |
| 31 | Gemma 4 31B IT | Google (Gemini / DeepMind) | 80.0% | DefaultThinking | Maker's own figureVendor-reportedGoogle ↗ | — | 2 Apr 2026 |
| 32 | Gemma 4 26B A4B IT | Google (Gemini / DeepMind) | 77.1% | DefaultThinking | Maker's own figureVendor-reportedGoogle ↗ | — | 2 Apr 2026 |
| 33 | Gemini 3.1 Flash-Lite | Google (Gemini / DeepMind) | 72.0% | High | Maker's own figureVendor-reportedGoogle ↗ | — | 3 Mar 2026 |
| 34 | Mistral Small 4 | Mistral AI | 63.6% | HighMistral Small 4 - High | Maker's own figureVendor-reportedMistral AI ↗ | — | 16 Mar 2026 |