TestsBenchmarks · Math
FrontierMath (Tiers 1-3)
Accuracy on hundreds of unpublished, expert-written research-level math problems (Epoch-run, v2).
% solvedHigher is betterToo easy now: the top models all score near perfectSaturatedThe test's websiteOfficial page ↗
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | GPT-6.1 Sol | OpenAI | 93.7% | Maxgpt-6.1-sol_max | Independent testIndependentEpoch AI ↗ | — | 29 Sep 2026 |
| 2 | GPT-6 Astra | OpenAI | 93.7% | Maxgpt-6-astra_max | Independent testIndependentEpoch AI ↗ | — | 30 Aug 2026 |
| 3 | Claude Opus 5.5 | Anthropic | 91.2% | Maxclaude-opus-5-5_max | Independent testIndependentEpoch AI ↗ | — | 22 Sep 2026 |
| 4 | Claude Fable 5.1 | Anthropic | 90.2% | Maxclaude-fable-5-1_max | Independent testIndependentEpoch AI ↗ | — | 1 Sep 2026 |
| 5 | GPT-6 Sol | OpenAI | 89.8% | Maxgpt-6-sol_max | Independent testIndependentEpoch AI ↗ | — | 22 Sep 2026 |
| 6 | GPT-5.6 Sol | OpenAI | 89.1% | Maxgpt-5.6-sol_max | Independent testIndependentEpoch AI ↗ | — | 9 Jul 2026 |
| 7 | Claude Sonnet 5.5 | Anthropic | 88.8% | Maxclaude-sonnet-5-5_max | Independent testIndependentEpoch AI ↗ | — | 29 Sep 2026 |
| 8 | GPT-5.5 Pro | OpenAI | 87.7% | Extra highgpt-5.5-pro_xhigh | Independent testIndependentEpoch AI ↗ | — | 12 Jun 2026 |
| 9 | Claude Fable 5 | Anthropic | 87.0% | Maxclaude-fable-5_max | Independent testIndependentEpoch AI ↗ | — | 9 Jun 2026 |
| 10 | GPT-5.6 Terra | OpenAI | 86.0% | Maxgpt-5.6-terra_max | Independent testIndependentEpoch AI ↗ | — | 9 Jul 2026 |
| 11 | Claude Opus 5 | Anthropic | 85.6% | Maxclaude-opus-5_max | Independent testIndependentEpoch AI ↗ | — | 24 Jul 2026 |
| 12 | GPT-5.5 | OpenAI | 85.3% | Extra highgpt-5.5_xhigh | Independent testIndependentEpoch AI ↗ | — | 11 Jun 2026 |
| 13 | GPT-5.4 Pro | OpenAI | 82.5% | Extra highgpt-5.4-pro-2026-03-05_xhigh | Independent testIndependentEpoch AI ↗ | — | 13 Jun 2026 |
| 14 | GPT-5.6 Luna | OpenAI | 82.1% | Maxgpt-5.6-luna_max | Independent testIndependentEpoch AI ↗ | — | 9 Jul 2026 |
| 15 | Claude Opus 4.8 | Anthropic | 80.0% | Maxclaude-opus-4-8_max | Independent testIndependentEpoch AI ↗ | — | 10 Jun 2026 |
| 16 | GPT-6 Luna | OpenAI | 78.9% | Maxgpt-6-luna_max | Independent testIndependentEpoch AI ↗ | — | 22 Sep 2026 |
| 17 | GPT-5.4 | OpenAI | 78.6% | Extra highgpt-5.4-2026-03-05_xhigh | Independent testIndependentEpoch AI ↗ | — | 11 Jun 2026 |
| 18 | Qwen3.8 Max (0803) | Qwen (Alibaba) | 74.7% | Extra highqwen3.8-max_xhigh | Independent testIndependentEpoch AI ↗ | — | 4 Aug 2026 |
| 19 | Muse Spark 1.3 | Meta (Meta Superintelligence Labs, Muse) | 74.4% | Extra highmuse-spark-1.3_xhigh | Independent testIndependentEpoch AI ↗ | — | 16 Sep 2026 |
| 20 | GPT-5.2 Pro | OpenAI | 74.0% | Extra highgpt-5.2-pro-2025-12-11_xhigh | Independent testIndependentEpoch AI ↗ | — | 13 Jun 2026 |
| 21 | Kimi K3 | Moonshot AI (Kimi) | 72.2% | Maxkimi-k3_max | Independent testIndependentEpoch AI ↗ | — | 17 Jul 2026 |
| 22 | Gemini 3.7 Flash | Google (Gemini / DeepMind) | 71.6% | Highgemini-3.7-flash_high | Independent testIndependentEpoch AI ↗ | — | 14 Aug 2026 |
| 23 | Claude Opus 4.7 | Anthropic | 70.2% | Maxclaude-opus-4-7_max | Independent testIndependentEpoch AI ↗ | — | 10 Jun 2026 |
| 24 | GLM-5.3 | Z.ai (Zhipu) | 68.8% | Maxglm-5.3_max | Independent testIndependentEpoch AI ↗ | — | 25 Aug 2026 |