TestsBenchmarks · Coding
FrontierSWE v2
Proximal's 34 ultra-long-horizon engineering/research tasks (~20h per task), run in Proximal's harness at max effort.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | OpenAI | 65.5% | Maxreasoning effort=max | Independent testIndependentProximal (via Anthropic system card) ↗ | Proximal harness | 22 Sep 2026 |
| 2 | Claude Opus 5.5 | Anthropic | 62.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Proximal harness | 22 Sep 2026 |
| 3 | Claude Sonnet 5.5 | Anthropic | 61.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Proximal harness | 28 Sep 2026 |
| 4 | Claude Fable 5.1 | Anthropic | 56.3% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Proximal harness | 22 Sep 2026 |
| 5 | Gemini 4 Argon | Google (Gemini / DeepMind) | 55.0% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | — | 30 Sep 2026 |
| 6 | Claude Opus 5 | Anthropic | 52.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Proximal harness | 1 Sep 2026 |
| 7 | Claude Fable 5 | Anthropic | 48.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Proximal harness | 1 Sep 2026 |
| 8 | GPT-5.6 Sol | OpenAI | 32.2% | Maxreasoning effort=max | Independent testIndependentProximal (via Anthropic system card) ↗ | Proximal harness | 22 Sep 2026 |