TestsBenchmarks · Knowledge
MMMLU
Multilingual MMLU.
% solvedHigher is betterToo easy now: the top models all score near perfectSaturatedThe test's websiteOfficial page ↗
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 92.6% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | — | 19 Feb 2026 |
| 2 | Qwen3.7 Max | Qwen (Alibaba) | 90.3% | Defaultthinking (vendor recommends 'Reasoning effort is set to xhigh' system prompt for reasoning tasks) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | — | 16 May 2026 |
| 3 | Qwen3.6 Plus | Qwen (Alibaba) | 89.5% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | — | 2 Apr 2026 |
| 4 | Qwen3.7 Plus | Qwen (Alibaba) | 89.0% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | — | 21 May 2026 |
| 5 | Gemini 3.1 Flash-Lite | Google (Gemini / DeepMind) | 88.9% | High | Maker's own figureVendor-reportedGoogle ↗ | — | 3 Mar 2026 |
| 6 | Gemma 4 31B IT | Google (Gemini / DeepMind) | 88.4% | DefaultThinking | Maker's own figureVendor-reportedGoogle ↗ | — | 2 Apr 2026 |
| 7 | Gemma 4 26B A4B IT | Google (Gemini / DeepMind) | 86.3% | DefaultThinking | Maker's own figureVendor-reportedGoogle ↗ | — | 2 Apr 2026 |