Agentic & Tool Use benchmark
τ²-bench Banking leaderboard
As of October 2026, Qwen3.8 Max has the highest published τ²-bench Banking score on Noometry at 55.1%, out of 26 models with results.
Last verified
About τ²-bench Banking
Banking customer-service tasks that need the agent to search a knowledge base as well as use account tools under policy.
- Category
- Agentic & Tool Use
- Format
- Tool use with retrieval
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- taubench.com
Top 15 models
- Qwen3.8 Max 55.1%
- Claude Opus 5 48.7%
- Grok 4.5 47.9%
- GPT-5.6 Sol 46.9%
- GPT-5.5 44.6%
- Muse Spark 1.1 40.5%
- Claude Opus 4.7 40.2%
- Claude Fable 5 39.7%
- Claude Opus 4.8 39.7%
- GPT-5.4 39.4%
- GLM-5.2 37.1%
- Kimi K3 37.1%
- GPT-5.2 32.2%
- Claude Opus 4.6 27.3%
- Gemini 3 Flash Preview 27.3%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does τ²-bench Banking measure?
Banking customer-service tasks that need the agent to search a knowledge base as well as use account tools under policy.
Which model has the highest τ²-bench Banking score?
As of October 2026, Qwen3.8 Max has the highest published τ²-bench Banking score on Noometry at 55.1%, out of 26 models with results.
What is the best open-weight model on τ²-bench Banking?
GLM-5.2 has the highest τ²-bench Banking accuracy among open-weight models at 37.1%, ranking 11 of 26 overall.