TestsBenchmarks · Knowledge
GDPval-AA (v1)
Artificial Analysis agentic GDPval variant, Elo from blind pairwise comparisons.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 1932 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 9 Jun 2026 |
| 2 | Claude Opus 4.8 | Anthropic | 1890 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 9 Jun 2026 |
| 3 | Claude Opus 4.7 | Anthropic | 1753 | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 28 May 2026 |
| 4 | Gemini 3.5 Flash | Google (Gemini / DeepMind) | 1656 | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | — | 19 May 2026 |
| 5 | DeepSeek V4 Pro (Preview, 0423) | DeepSeek | 1554 | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | — | 24 Apr 2026 |
| 6 | MiniMax M2.7 | MiniMax | 1495 | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | — | 18 Mar 2026 |
| 7 | DeepSeek V4 Flash (Preview, 0423) | DeepSeek | 1395 | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | — | 24 Apr 2026 |
| 8 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 1317 | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | — | 19 Feb 2026 |
| 9 | Muse Glimmer 30B | Meta (Meta Superintelligence Labs, Muse) | 953 | HighHigh reasoning | Maker's own figureVendor-reportedMeta ↗ | — | 10 Aug 2026 |