TestsBenchmarks · Coding
GSO
Share of 102 software-optimization tasks where one attempt reaches at least 95% of the expert speedup while passing correctness tests.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | Anthropic | 88.2% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | OpenHands | 27 Sep 2026 |
| 2 | GPT-6 Astra | OpenAI | 79.4% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | OpenHands | 27 Sep 2026 |
| 3 | Claude Fable 5 | Anthropic | 78.4% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | OpenHands | 27 Sep 2026 |
| 4 | GPT-5.6 Sol | OpenAI | 76.5% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | OpenHands | 27 Sep 2026 |
| 5 | Claude Opus 4.8 | Anthropic | 47.1% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | OpenHands | 12 Jul 2026 |
| 6 | Claude Opus 4.7 | Anthropic | 44.1% | Highreasoning_effort=high | Independent testIndependentGSO ↗ | OpenHands | 27 Apr 2026 |
| 7 | Claude Opus 4.6 | Anthropic | 41.2% | Highreasoning_effort=high | Independent testIndependentGSO ↗ | OpenHands | 27 Apr 2026 |
| 8 | GPT-5.5 | OpenAI | 40.2% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | OpenHands | 27 Apr 2026 |
| 9 | Claude Sonnet 5 | Anthropic | 37.3% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | OpenHands | 12 Jul 2026 |
| 10 | GPT-5.4 | OpenAI | 31.4% | Extra highreasoning_effort=xhigh | Independent testIndependentGSO ↗ | OpenHands | 10 Mar 2026 |
| 11 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 22.6% | Default | Independent testIndependentGSO ↗ | OpenHands | 9 Mar 2026 |