Agentic & Tool Use benchmark
Terminal-Bench leaderboard
As of October 2026, GPT-5.5 has the highest published Terminal-Bench score on Noometry at 84.7%, out of 41 models with results.
Last verified
About Terminal-Bench
Hard tasks completed in a real terminal: compiling code, configuring systems, training models and debugging environments.
- Category
- Agentic & Tool Use
- Introduced
- 2025
- Format
- Agent in a shell
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- www.tbench.ai
Top 15 models
- GPT-5.5 84.7%
- GPT-5.4 81.8%
- Claude Opus 4.7 80.2%
- Gemini 3.1 Pro Preview 80.2%
- Claude Opus 4.6 79.8%
- GPT-5.3 Codex 78.4%
- Gemini 3 Pro 69.4%
- GPT-5.2 Codex 66.5%
- GPT-5.2 64.9%
- Gemini 3 Flash Preview 64.3%
- Claude Opus 4.5 63.1%
- GPT-5.1-Codex-mini 61.6%
- GPT-5.1-Codex 60.4%
- Grok 4.20 (Non-Reasoning) 57.3%
- Claude Sonnet 4.6 53.4%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does Terminal-Bench measure?
Hard tasks completed in a real terminal: compiling code, configuring systems, training models and debugging environments.
Which model has the highest Terminal-Bench score?
As of October 2026, GPT-5.5 has the highest published Terminal-Bench score on Noometry at 84.7%, out of 41 models with results.
What is the best open-weight model on Terminal-Bench?
GLM-5 has the highest Terminal-Bench accuracy among open-weight models at 52.4%, ranking 16 of 41 overall.