TestsBenchmarks · Coding
SWE-bench Pro (Anthropic internal subset)
Anthropic-internal 478-problem subset of SWE-bench Pro used for cost/effort studies; not comparable to the public leaderboard.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | Anthropic | 95.3% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | — | 1 Oct 2026 |
| 2 | Claude Fable 5.1 | Anthropic | 92.3% | Highadaptive thinking, effort=high (default) | Maker's own figureVendor-reportedAnthropic ↗ | — | 1 Oct 2026 |
| 3 | Claude Sonnet 5 | Anthropic | 77.4% | Highadaptive thinking, effort=high (default) | Maker's own figureVendor-reportedAnthropic ↗ | — | 1 Oct 2026 |