TestsBenchmarks · Reasoning
ARC-AGI-1
Abstract visual reasoning puzzles (v1).
% solvedHigher is betterToo easy now: the top models all score near perfectSaturatedThe test's websiteOfficial page ↗
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 98.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 1 Sep 2026 |
| 2 | GPT-6 Astra | OpenAI | 98.5% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | — | 3 Sep 2026 |
| 3 | Claude Opus 5 | Anthropic | 97.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 24 Jul 2026 |
| 4 | Claude Fable 5.1 | Anthropic | 97.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 1 Sep 2026 |
| 5 | GPT-5.6 Sol | OpenAI | 97.5% | Defaultbest score across efforts (OpenAI table: "maximum at any effort") | Maker's own figureVendor-reportedOpenAI ↗ | — | 3 Sep 2026 |
| 6 | GPT-5.5 | OpenAI | 95.0% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | — | 23 Apr 2026 |
| 7 | Claude Opus 4.8 | Anthropic | 92.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 24 Jul 2026 |