Coding benchmark

SWE-bench Verified (bash only) leaderboard

As of October 2026, Claude Opus 4.5 has the highest published SWE-bench Verified (bash only) score on Noometry at 76.8%, out of 39 models with results.

Last verified

About SWE-bench Verified (bash only)

The 500 human-validated SWE-bench Verified GitHub issues, solved by every model inside the same minimal bash-only agent, so the score reflects the model rather than the scaffold.

Category
Coding
Introduced
2025
Size
500 tasks
Format
Repository patch
Unit
Percent (random guessing ≈ 0%)
Official site
www.swebench.com

Top 15 models

Top models on SWE-bench Verified (bash only)
  1. Claude Opus 4.5 76.8%
  2. Gemini 3 Flash Preview 75.8%
  3. MiniMax-M2.5 75.8%
  4. Claude Opus 4.6 75.6%
  5. Gemini 3 Pro 74.2%
  6. GLM-5 72.8%
  7. GPT-5.2 72.8%
  8. GPT-5.2 Codex 72.8%
  9. Claude Sonnet 4.5 71.4%
  10. Kimi K2.5 70.8%
  11. DeepSeek-V3.2-Exp 70%
  12. Claude Opus 4 67.6%
  13. Claude Haiku 4.5 66.6%
  14. GPT-5.1 66%
  15. GPT-5.1-Codex 66%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

SWE-bench Verified (bash only) results by model
#ModelProviderScoreSettingSourceDate
1Claude Opus 4.5 Anthropic76.8%highSWE-bench2026-02-17
2Gemini 3 Flash Preview Google75.8%highSWE-bench2026-02-17
3MiniMax-M2.5 MiniMax75.8%highSWE-bench2026-02-17
4Claude Opus 4.6 Anthropic75.6%SWE-bench2026-02-17
5Gemini 3 Pro Google74.2%SWE-bench2025-11-18
6GLM-5 Z.ai (Zhipu)72.8%highSWE-bench2026-02-17
7GPT-5.2 OpenAI72.8%highSWE-bench2026-02-17
8GPT-5.2 Codex OpenAI72.8%SWE-bench2026-02-19
9Claude Sonnet 4.5 Anthropic71.4%highSWE-bench2026-02-17
10Kimi K2.5 Moonshot AI70.8%highSWE-bench2026-02-17
11DeepSeek-V3.2-Exp DeepSeek70%highSWE-bench2026-02-17
12Claude Opus 4 Anthropic67.6%SWE-bench2025-08-02
13Claude Haiku 4.5 Anthropic66.6%highSWE-bench2026-02-17
14GPT-5.1 OpenAI66%mediumSWE-bench2025-11-20
15GPT-5.1-Codex OpenAI66%mediumSWE-bench2025-11-24
16GPT-5 OpenAI65%mediumSWE-bench2025-08-07
17Claude Sonnet 4 Anthropic64.9%SWE-bench2025-07-26
18Kimi K2 (Jul 2025) Moonshot AI63.4%SWE-bench2025-12-10
19MiniMax-M2 MiniMax61%SWE-bench2025-11-24
20GPT-5 Mini OpenAI59.8%mediumSWE-bench2025-08-07
21o3 OpenAI58.4%SWE-bench2025-07-26
22Devstral Small 2505 Mistral AI56.4%SWE-bench2025-12-09
23GLM-4.6 Z.ai (Zhipu)55.4%SWE-bench2025-12-01
24Qwen3-Coder 480B-A35B Instruct Alibaba (Qwen)55.4%SWE-bench2025-08-02
25GLM-4.5 Z.ai (Zhipu)54.2%SWE-bench2025-08-22
26Devstral 2 Mistral AI53.8%SWE-bench2025-12-09
27Gemini 2.5 Pro Google53.6%SWE-bench2025-07-26
28Claude 3.7 Sonnet Anthropic52.8%SWE-bench2025-07-20
29o4-mini OpenAI45%SWE-bench2025-07-26
30GPT-4.1 OpenAI39.6%SWE-bench2025-07-26
31GPT-5 Nano OpenAI34.8%mediumSWE-bench2025-08-07
32Gemini 2.5 Flash Google28.7%SWE-bench2025-07-26
33gpt-oss-120b OpenAI26%SWE-bench2025-08-07
34GPT-4.1 mini OpenAI23.9%SWE-bench2025-07-20
35GPT-4o OpenAI21.6%SWE-bench2025-07-20
36Llama 4 Maverick Meta21%SWE-bench2025-07-20
37Gemini 2.0 Flash (Feb 2025) Google13.5%SWE-bench2025-07-26
38Llama 4 Scout Meta9.1%SWE-bench2025-07-20
39Qwen2.5-Coder-32B Alibaba (Qwen)9%SWE-bench2025-08-03

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does SWE-bench Verified (bash only) measure?

The 500 human-validated SWE-bench Verified GitHub issues, solved by every model inside the same minimal bash-only agent, so the score reflects the model rather than the scaffold.

Which model has the highest SWE-bench Verified (bash only) score?

As of October 2026, Claude Opus 4.5 has the highest published SWE-bench Verified (bash only) score on Noometry at 76.8%, out of 39 models with results.

What is the best open-weight model on SWE-bench Verified (bash only)?

MiniMax-M2.5 has the highest SWE-bench Verified (bash only) accuracy among open-weight models at 75.8%, ranking 3 of 39 overall.