Skip to content
Bencher

TestsBenchmarks · Tool use

tau2-bench

Success rate of agents handling simulated customer-service conversations with tools and policies.

% solvedHigher is betterThe test's websiteOfficial page ↗

The best 11 · 11 models tested

Top 11 · best setting per model · 11 models, 19 results (17 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
tau2-bench leaderboard
Thinking levelSetting
1GLM-5.2Z.ai (Zhipu)99.1%MaxGLM-5.2 (Max)Independent testIndependentArtificial Analysis ↗—1 Oct 2026
2Claude Fable 5Anthropic98.5%MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Independent testIndependentArtificial Analysis ↗—1 Oct 2026
3Gemini 3.5 FlashGoogle (Gemini / DeepMind)95.6%MediumGemini 3.5 Flash (Medium)Independent testIndependentArtificial Analysis ↗—1 Oct 2026
4Claude Opus 4.8Anthropic94.4%MaxClaude Opus 4.8 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗—1 Oct 2026
5GPT-5.5OpenAI93.9%Extra highGPT-5.5 (Xhigh)Independent testIndependentArtificial Analysis ↗—1 Oct 2026
6Claude Opus 4.7Anthropic88.6%MaxClaude Opus 4.7 (Adaptive Reasoning, Max Effort)Independent testIndependentArtificial Analysis ↗—1 Oct 2026
7GPT-5.4OpenAI87.1%Extra highGPT-5.4 (Xhigh)Independent testIndependentArtificial Analysis ↗—1 Oct 2026
8GPT-5.6 TerraOpenAI86.3%MaxGPT-5.6 Terra (Max)Independent testIndependentArtificial Analysis ↗—1 Oct 2026
9GPT-5.6 SolOpenAI85.1%MaxGPT-5.6 Sol (Max)Independent testIndependentArtificial Analysis ↗—1 Oct 2026
10Gemma 4 31B ITGoogle (Gemini / DeepMind)76.9%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗—2 Apr 2026
11Gemma 4 26B A4B ITGoogle (Gemini / DeepMind)68.2%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗—2 Apr 2026