TestsBenchmarks · Agentic
Terminal-Bench 3.0
Share of 74 frontier-difficulty terminal tasks (v0.1 rolling set, formerly Frontier-Bench) an agent completes; superseded by 4.0.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 42.7% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | mini-SWE-agent | 1 Oct 2026 |
| 2 | GPT-5.6 Sol | OpenAI | 34.6% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Codex | 1 Oct 2026 |
| 3 | Claude Fable 5 | Anthropic | 34.1% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Claude Code | 1 Oct 2026 |
| 4 | GLM-5.3 | Z.ai (Zhipu) | 32.4% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Claude Code | 1 Oct 2026 |
| 5 | DeepSeek V4.1 Flash | DeepSeek | 30.0% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | DeepSeek Harness (Minimal) | 10 Sep 2026 |
| 6 | Grok 4.6 | SpaceXAI (formerly xAI) | 26.5% | High | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Grok Build | 1 Oct 2026 |
| 7 | Claude Opus 4.8 | Anthropic | 21.1% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Claude Code | 1 Oct 2026 |
| 8 | GPT-5.6 Terra | OpenAI | 20.8% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Codex | 1 Oct 2026 |
| 9 | SWE-1.7 Lightning | Cognition | 18.6% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Devin | 1 Oct 2026 |
| 10 | Grok 4.5 | SpaceXAI (formerly xAI) | 15.7% | Extra high | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Cursor CLI | 1 Oct 2026 |
| 11 | Gemini 3.7 Flash | Google (Gemini / DeepMind) | 14.9% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | mini-swe-agent (LiteLLM 1.96) | 13 Aug 2026 |
| 12 | Claude Sonnet 5 | Anthropic | 14.6% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Claude Code | 1 Oct 2026 |
| 13 | GPT-5.6 Luna | OpenAI | 14.3% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Codex | 1 Oct 2026 |
| 14 | DeepSeek V4 Pro (0813) | DeepSeek | 11.8% | MaxMax reasoning effort | Maker's own figureVendor-reportedDeepSeek ↗ | DeepSeek Harness (minimal mode) | 10 Sep 2026 |
| 15 | DeepSeek V4 Flash (0731) | DeepSeek | 7.6% | MaxMax reasoning effort | Maker's own figureVendor-reportedDeepSeek ↗ | DeepSeek Harness (minimal mode) | 10 Sep 2026 |
| 16 | Claude Opus 4.7 | Anthropic | 7.0% | Highshipped default effort (high) | Maker's own figureVendor-reportedAnthropic ↗ | Claude Managed Agents | 1 Oct 2026 |
| 17 | Gemini 3.6 Flash | Google (Gemini / DeepMind) | 5.4% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | mini-swe-agent | 13 Aug 2026 |
| 18 | GLM-5.2 | Z.ai (Zhipu) | 4.6% | Max | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | Claude Code | 1 Oct 2026 |