TestsBenchmarks · Tool use
tau2-bench
Success rate of agents handling simulated customer-service conversations with tools and policies.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | GLM-5.2 | Z.ai (Zhipu) | 99.1% | MaxGLM-5.2 (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 2 | Claude Fable 5 | Anthropic | 98.5% | MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 3 | Gemini 3.5 Flash | Google (Gemini / DeepMind) | 95.6% | MediumGemini 3.5 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 4 | Claude Opus 4.8 | Anthropic | 94.4% | MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 5 | GPT-5.5 | OpenAI | 93.9% | Extra highGPT-5.5 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 6 | Claude Opus 4.7 | Anthropic | 88.6% | MaxClaude Opus 4.7 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 7 | GPT-5.4 | OpenAI | 87.1% | Extra highGPT-5.4 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 8 | GPT-5.6 Terra | OpenAI | 86.3% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 9 | GPT-5.6 Sol | OpenAI | 85.1% | MaxGPT-5.6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 10 | Gemma 4 31B IT | Google (Gemini / DeepMind) | 76.9% | DefaultThinking | Maker's own figureVendor-reportedGoogle ↗ | — | 2 Apr 2026 |
| 11 | Gemma 4 26B A4B IT | Google (Gemini / DeepMind) | 68.2% | DefaultThinking | Maker's own figureVendor-reportedGoogle ↗ | — | 2 Apr 2026 |