TestsBenchmarks · Math
AIME 2026
2026 American Invitational Mathematics Examination, no tools.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | GLM-5.2 | Z.ai (Zhipu) | 99.2% | Defaultthinking (effort not stated) | Maker's own figureVendor-reportedZ.ai ↗ | — | 16 Jun 2026 |
| 2 | Kimi K2.6 | Moonshot AI (Kimi) | 96.4% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | — | 20 Apr 2026 |
| 3 | Qwen3.6 Plus | Qwen (Alibaba) | 95.3% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | — | 2 Apr 2026 |
| 4 | GLM-5.1 | Z.ai (Zhipu) | 95.3% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | — | 7 Apr 2026 |
| 5 | Muse Glimmer 30B | Meta (Meta Superintelligence Labs, Muse) | 94.7% | HighHigh reasoning | Maker's own figureVendor-reportedMeta ↗ | — | 10 Aug 2026 |
| 6 | Qwen3.6 27B | Qwen (Alibaba) | 94.1% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | — | 22 Apr 2026 |
| 7 | Gemma 4 31B IT | Google (Gemini / DeepMind) | 89.2% | DefaultThinking | Maker's own figureVendor-reportedGoogle ↗ | — | 2 Apr 2026 |
| 8 | Gemma 4 26B A4B IT | Google (Gemini / DeepMind) | 88.3% | DefaultThinking | Maker's own figureVendor-reportedGoogle ↗ | — | 2 Apr 2026 |