TestsBenchmarks · Coding
CursorBench 3.2.0
Earlier CursorBench release; not comparable to 4.0.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | Anthropic | 73.4% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Cursor agent | 1 Sep 2026 |
| 2 | Claude Fable 5 | Anthropic | 70.5% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Cursor agent | 1 Sep 2026 |
| 3 | Claude Opus 5 | Anthropic | 70.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Cursor agent | 1 Sep 2026 |
| 4 | Grok 4.6 | SpaceXAI (formerly xAI) | 69.9% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | — | 12 Aug 2026 |
| 5 | GPT-5.6 Sol | OpenAI | 67.2% | Maxreasoning effort=max | Independent testIndependentCursor (via Anthropic launch post) ↗ | Cursor agent | 1 Sep 2026 |
| 6 | Grok 4.5 | SpaceXAI (formerly xAI) | 66.7% | High | Maker's own figureVendor-reportedSpaceXAI ↗ | — | 12 Aug 2026 |