TestsBenchmarks · Agentic
Frontier-Bench v0.1
Agentic terminal-coding benchmark (74 tasks) run via Harbor or mini-SWE-agent; reported at Claude Opus 5 launch.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 44.4% | Extra highadaptive thinking, effort=xhigh | Maker's own figureVendor-reportedAnthropic ↗ | mini-SWE-agent (GKE) | 24 Jul 2026 |
| 2 | Claude Fable 5 | Anthropic | 33.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 24 Jul 2026 |
| 3 | Claude Opus 4.8 | Anthropic | 21.1% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 24 Jul 2026 |
| 4 | Claude Sonnet 5 | Anthropic | 17.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | mini-SWE-agent (GKE) | 24 Jul 2026 |