TestsBenchmarks · Agentic
Terminal-Bench-Science 0.1
70 scientific research workflows (data analysis, simulation, theorem proving) solved by an agent CLI and graded by hidden tests.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | OpenAI | 68.1% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | Codex | 29 Sep 2026 |
| 2 | Claude Opus 5.5 | Anthropic | 63.3% | Maxeffort=max | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | — | 29 Sep 2026 |
| 3 | Claude Sonnet 5.5 | Anthropic | 59.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Claude Code (--bare) | 28 Sep 2026 |
| 4 | Gemini 4 Argon | Google (Gemini / DeepMind) | 57.6% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | — | 30 Sep 2026 |
| 5 | GPT-6.1 Sol | OpenAI | 57.0% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | Codex | 29 Sep 2026 |
| 6 | Claude Fable 5.1 | Anthropic | 52.6% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Claude Code (--bare) | 1 Sep 2026 |
| 7 | Claude Opus 5 | Anthropic | 29.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | Claude Code (--bare) | 22 Sep 2026 |
| 8 | GPT-6 Sol | OpenAI | 27.6% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | Codex | 29 Sep 2026 |
| 9 | Claude Fable 5 | Anthropic | 25.0% | Highadaptive thinking, effort=high | Maker's own figureVendor-reportedAnthropic ↗ | Claude Code (--bare) | 1 Sep 2026 |
| 10 | GPT-5.6 Sol | OpenAI | 22.4% | Maxreasoning effort=max | Maker's own figureVendor-reportedOpenAI ↗ | Codex | 3 Sep 2026 |