Skip to content
Bencher

TestsBenchmarks · Agentic

Terminal-Bench 3.0

Share of 74 frontier-difficulty terminal tasks (v0.1 rolling set, formerly Frontier-Bench) an agent completes; superseded by 4.0.

% solvedHigher is betterCounts toward:Weights: Coding ×0.5The test's websiteOfficial page ↗

The best 15 · 18 models tested

Top 15 · best setting per model · 18 models, 21 results (12 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
Terminal-Bench 3.0 leaderboard
Thinking levelSetting
1Claude Opus 5Anthropic42.7%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗mini-SWE-agent1 Oct 2026
2GPT-5.6 SolOpenAI34.6%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗Codex1 Oct 2026
3Claude Fable 5Anthropic34.1%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗Claude Code1 Oct 2026
4GLM-5.3Z.ai (Zhipu)32.4%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗Claude Code1 Oct 2026
5DeepSeek V4.1 FlashDeepSeek30.0%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗DeepSeek Harness (Minimal)10 Sep 2026
6Grok 4.6SpaceXAI (formerly xAI)26.5%HighIndependent testIndependentSnorkel AI / Terminal-Bench ↗Grok Build1 Oct 2026
7Claude Opus 4.8Anthropic21.1%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗Claude Code1 Oct 2026
8GPT-5.6 TerraOpenAI20.8%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗Codex1 Oct 2026
9SWE-1.7 LightningCognition18.6%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗Devin1 Oct 2026
10Grok 4.5SpaceXAI (formerly xAI)15.7%Extra highIndependent testIndependentSnorkel AI / Terminal-Bench ↗Cursor CLI1 Oct 2026
11Gemini 3.7 FlashGoogle (Gemini / DeepMind)14.9%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗mini-swe-agent (LiteLLM 1.96)13 Aug 2026
12Claude Sonnet 5Anthropic14.6%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗Claude Code1 Oct 2026
13GPT-5.6 LunaOpenAI14.3%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗Codex1 Oct 2026
14DeepSeek V4 Pro (0813)DeepSeek11.8%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗DeepSeek Harness (minimal mode)10 Sep 2026
15DeepSeek V4 Flash (0731)DeepSeek7.6%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗DeepSeek Harness (minimal mode)10 Sep 2026
16Claude Opus 4.7Anthropic7.0%Highshipped default effort (high)Maker's own figureVendor-reportedAnthropic ↗Claude Managed Agents1 Oct 2026
17Gemini 3.6 FlashGoogle (Gemini / DeepMind)5.4%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗mini-swe-agent13 Aug 2026
18GLM-5.2Z.ai (Zhipu)4.6%MaxIndependent testIndependentSnorkel AI / Terminal-Bench ↗Claude Code1 Oct 2026