TestsBenchmarks · Agentic
METR 50% time horizon
Length of software tasks (in human-expert minutes) an agent completes with 50% reliability.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Mythos Preview | Anthropic | 17.4 h | Default | Independent testIndependentMETR ↗ | METR react agent (Inspect) | 8 May 2026 |
| 2 | Claude Opus 4.6 | Anthropic | 12.0 h | Default | Independent testIndependentMETR ↗ | METR react agent (Inspect) | 8 May 2026 |
| 3 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 6.4 h | Default | Independent testIndependentMETR ↗ | METR react agent (Inspect) | 8 May 2026 |
| 4 | GPT-5.2 | OpenAI | 5.9 h | Default | Independent testIndependentMETR ↗ | METR react agent (Inspect) | 8 May 2026 |
| 5 | GPT-5.3-Codex | OpenAI | 5.8 h | Default | Independent testIndependentMETR ↗ | METR react agent (Inspect) | 8 May 2026 |
| 6 | GPT-5.4 | OpenAI | 5.7 h | Default | Independent testIndependentMETR ↗ | METR react agent (Inspect) | 8 May 2026 |
| 7 | Claude Opus 4.5 | Anthropic | 4.9 h | Default | Independent testIndependentMETR ↗ | METR react agent (Inspect) | 8 May 2026 |
| 8 | Gemini 3 Pro | Google (Gemini / DeepMind) | 3.7 h | Default | Independent testIndependentMETR ↗ | METR react agent (Inspect) | 8 May 2026 |
| 9 | GPT-5.1-Codex-Max | OpenAI | 3.7 h | Default | Independent testIndependentMETR ↗ | METR react agent (Inspect) | 8 May 2026 |