Agentic & Tool Use benchmark
τ²-bench Retail leaderboard
As of October 2026, Qwen3.5 397B-A17B has the highest published τ²-bench Retail score on Noometry at 84.4%, out of 7 models with results.
Last verified
About τ²-bench Retail
Retail customer-service conversations (returns, exchanges, order changes) handled with tools under a written policy.
- Category
- Agentic & Tool Use
- Introduced
- 2025
- Format
- Tool use with a simulated user
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- taubench.com
Top 7 models
- Qwen3.5 397B-A17B 84.4%
- GPT-5.2 81.6%
- Claude Opus 4.5 79.6%
- Gemini 3 Flash Preview 76.8%
- Gemini 3 Pro 75.9%
- GLM-5 73.7%
- Claude Sonnet 4.5 72.4%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | Qwen3.5 397B-A17B | 84.4% | enabled | τ²-bench | 2026-03-02 | |
| 2 | GPT-5.2 | OpenAI | 81.6% | high | τ²-bench | 2026-02-26 |
| 3 | Claude Opus 4.5 | Anthropic | 79.6% | high | τ²-bench | 2026-02-26 |
| 4 | Gemini 3 Flash Preview | 76.8% | high | τ²-bench | 2026-03-02 | |
| 5 | Gemini 3 Pro | 75.9% | high | τ²-bench | 2026-03-02 | |
| 6 | GLM-5 | Z.ai (Zhipu) | 73.7% | enabled | τ²-bench | 2026-03-02 |
| 7 | Claude Sonnet 4.5 | Anthropic | 72.4% | enabled | τ²-bench | 2026-02-26 |
Compare the leaders
Frequently asked questions
What does τ²-bench Retail measure?
Retail customer-service conversations (returns, exchanges, order changes) handled with tools under a written policy.
Which model has the highest τ²-bench Retail score?
As of October 2026, Qwen3.5 397B-A17B has the highest published τ²-bench Retail score on Noometry at 84.4%, out of 7 models with results.
What is the best open-weight model on τ²-bench Retail?
Qwen3.5 397B-A17B has the highest τ²-bench Retail accuracy among open-weight models at 84.4%, ranking 1 of 7 overall.