Skip to content
Bencher

TestsBenchmarks · Agentic

METR 50% time horizon

Length of software tasks (in human-expert minutes) an agent completes with 50% reliability.

Time horizonHigher is betterThe test's websiteOfficial page ↗

The best 9 · 9 models tested

Top 9 · best setting per model · 9 models, 9 results (9 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
METR 50% time horizon leaderboard
Thinking levelSetting
1Claude Mythos PreviewAnthropic17.4 hDefaultIndependent testIndependentMETR ↗METR react agent (Inspect)8 May 2026
2Claude Opus 4.6Anthropic12.0 hDefaultIndependent testIndependentMETR ↗METR react agent (Inspect)8 May 2026
3Gemini 3.1 Pro (Preview)Google (Gemini / DeepMind)6.4 hDefaultIndependent testIndependentMETR ↗METR react agent (Inspect)8 May 2026
4GPT-5.2OpenAI5.9 hDefaultIndependent testIndependentMETR ↗METR react agent (Inspect)8 May 2026
5GPT-5.3-CodexOpenAI5.8 hDefaultIndependent testIndependentMETR ↗METR react agent (Inspect)8 May 2026
6GPT-5.4OpenAI5.7 hDefaultIndependent testIndependentMETR ↗METR react agent (Inspect)8 May 2026
7Claude Opus 4.5Anthropic4.9 hDefaultIndependent testIndependentMETR ↗METR react agent (Inspect)8 May 2026
8Gemini 3 ProGoogle (Gemini / DeepMind)3.7 hDefaultIndependent testIndependentMETR ↗METR react agent (Inspect)8 May 2026
9GPT-5.1-Codex-MaxOpenAI3.7 hDefaultIndependent testIndependentMETR ↗METR react agent (Inspect)8 May 2026