TestsBenchmarks · Multimodal
Blueprint-Bench 2
Spatial reasoning from floor plans/blueprints.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 38.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 9 Jun 2026 |
| 2 | Gemini 3.5 Flash | Google (Gemini / DeepMind) | 33.6% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | — | 19 May 2026 |
| 3 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 26.5% | Defaultthinking level not stated (API default high) | Maker's own figureVendor-reportedGoogle ↗ | — | 19 May 2026 |
| 4 | Claude Opus 4.8 | Anthropic | 14.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 9 Jun 2026 |