Agentic & Tool Use benchmark

Terminal-Bench leaderboard

As of October 2026, GPT-5.5 has the highest published Terminal-Bench score on Noometry at 84.7%, out of 41 models with results.

Last verified

About Terminal-Bench

Hard tasks completed in a real terminal: compiling code, configuring systems, training models and debugging environments.

Category
Agentic & Tool Use
Introduced
2025
Format
Agent in a shell
Unit
Percent (random guessing ≈ 0%)
Official site
www.tbench.ai

Top 15 models

Top models on Terminal-Bench
  1. GPT-5.5 84.7%
  2. GPT-5.4 81.8%
  3. Claude Opus 4.7 80.2%
  4. Gemini 3.1 Pro Preview 80.2%
  5. Claude Opus 4.6 79.8%
  6. GPT-5.3 Codex 78.4%
  7. Gemini 3 Pro 69.4%
  8. GPT-5.2 Codex 66.5%
  9. GPT-5.2 64.9%
  10. Gemini 3 Flash Preview 64.3%
  11. Claude Opus 4.5 63.1%
  12. GPT-5.1-Codex-mini 61.6%
  13. GPT-5.1-Codex 60.4%
  14. Grok 4.20 (Non-Reasoning) 57.3%
  15. Claude Sonnet 4.6 53.4%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Terminal-Bench results by model
#ModelProviderScoreSettingSourceDate
1GPT-5.5 OpenAI84.7%Epoch AI
2GPT-5.4 OpenAI81.8%Epoch AI
3Claude Opus 4.7 Anthropic80.2%Epoch AI
4Gemini 3.1 Pro Preview Google80.2%Epoch AI
5Claude Opus 4.6 Anthropic79.8%Epoch AI
6GPT-5.3 Codex OpenAI78.4%Epoch AI
7Gemini 3 Pro Google69.4%Epoch AI
8GPT-5.2 Codex OpenAI66.5%Epoch AI
9GPT-5.2 OpenAI64.9%Epoch AI
10Gemini 3 Flash Preview Google64.3%Epoch AI
11Claude Opus 4.5 Anthropic63.1%Epoch AI
12GPT-5.1-Codex-mini OpenAI61.6%Epoch AI
13GPT-5.1-Codex OpenAI60.4%Epoch AI
14Grok 4.20 (Non-Reasoning) xAI57.3%Epoch AI
15Claude Sonnet 4.6 Anthropic53.4%Epoch AI
16GLM-5 Z.ai (Zhipu)52.4%Epoch AI
17GPT-5 OpenAI49.6%Epoch AI
18GPT-5.1 OpenAI47.6%mediumEpoch AI
19Claude Sonnet 4.5 Anthropic46.5%Epoch AI
20MiniMax-M2.7 MiniMax45.1%Epoch AI
21GPT-5-Codex OpenAI44.3%Epoch AI
22Kimi K2.5 Moonshot AI43.2%Epoch AI
23MiniMax-M2.5 MiniMax42.7%Epoch AI
24DeepSeek-V3.2-Exp DeepSeek39.6%Epoch AI
25Claude Opus 4.1 Anthropic38%Epoch AI
26MiniMax-M2.1 MiniMax36.6%Epoch AI
27Kimi K2 (Jul 2025) Moonshot AI35.7%Epoch AI
28Claude Haiku 4.5 Anthropic35.5%Epoch AI
29GPT-5 Mini OpenAI34.8%Epoch AI
30GLM-4.7 Z.ai (Zhipu)33.4%Epoch AI
31Gemini 2.5 Pro Google32.6%Epoch AI
32MiniMax-M2 MiniMax30%Epoch AI
33Grok 4 xAI27.2%Epoch AI
34Qwen3-Coder 480B-A35B Instruct Alibaba (Qwen)27.2%Epoch AI
35GLM-4.6 Z.ai (Zhipu)24.5%Epoch AI
36Qwen3.6 35B-A3B Alibaba (Qwen)23%Epoch AI
37GPT-5 Nano OpenAI21.8%Epoch AI
38gpt-oss-120b OpenAI18.7%Epoch AI
39Gemini 2.5 Flash Google17.1%Epoch AI
40Qwen3.5-9B Alibaba (Qwen)9.2%Epoch AI
41gpt-oss-20b OpenAI3.4%Epoch AI

Compare the leaders

Other agentic & tool use benchmarks

Frequently asked questions

What does Terminal-Bench measure?

Hard tasks completed in a real terminal: compiling code, configuring systems, training models and debugging environments.

Which model has the highest Terminal-Bench score?

As of October 2026, GPT-5.5 has the highest published Terminal-Bench score on Noometry at 84.7%, out of 41 models with results.

What is the best open-weight model on Terminal-Bench?

GLM-5 has the highest Terminal-Bench accuracy among open-weight models at 52.4%, ranking 16 of 41 overall.