TestsBenchmarks · Tool use
tau3-bench
Third-generation tau-bench customer-service agent simulations with policy following.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.6 Plus | Qwen (Alibaba) | 70.7% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | — | 2 Apr 2026 |
| 2 | GLM-5.1 | Z.ai (Zhipu) | 70.6% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | — | 7 Apr 2026 |