Coding benchmark

SWE-bench Multilingual leaderboard

As of October 2026, Gemini 3 Flash Preview has the highest published SWE-bench Multilingual score on Noometry at 72.7%, out of 13 models with results.

Last verified

About SWE-bench Multilingual

Real GitHub issues from repositories in nine programming languages other than Python, solved in the same bash-only agent.

Category
Coding
Introduced
2025
Size
300 tasks
Format
Repository patch
Unit
Percent (random guessing ≈ 0%)
Official site
www.swebench.com

Top 13 models

Top models on SWE-bench Multilingual
  1. Gemini 3 Flash Preview 72.7%
  2. Claude Opus 4.6 72%
  3. Claude Opus 4.5 70.7%
  4. GLM-5 69.7%
  5. Gemini 3 Pro 68.7%
  6. MiniMax-M2.5 68.3%
  7. Kimi K2.5 67.3%
  8. Claude Sonnet 4.5 67%
  9. GPT-5.2 66.7%
  10. GPT-5.2 Codex 66.3%
  11. Claude Haiku 4.5 64.7%
  12. DeepSeek-V3.2-Exp 59%
  13. GPT-5 Mini 39.7%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

SWE-bench Multilingual results by model
#ModelProviderScoreSettingSourceDate
1Gemini 3 Flash Preview Google72.7%SWE-bench2026-02-13
2Claude Opus 4.6 Anthropic72%SWE-bench2026-02-13
3Claude Opus 4.5 Anthropic70.7%SWE-bench2026-02-13
4GLM-5 Z.ai (Zhipu)69.7%SWE-bench2026-02-13
5Gemini 3 Pro Google68.7%SWE-bench2026-02-13
6MiniMax-M2.5 MiniMax68.3%SWE-bench2026-02-16
7Kimi K2.5 Moonshot AI67.3%SWE-bench2026-02-13
8Claude Sonnet 4.5 Anthropic67%SWE-bench2026-02-13
9GPT-5.2 OpenAI66.7%highSWE-bench2026-02-13
10GPT-5.2 Codex OpenAI66.3%SWE-bench2026-02-20
11Claude Haiku 4.5 Anthropic64.7%SWE-bench2026-02-13
12DeepSeek-V3.2-Exp DeepSeek59%SWE-bench2026-02-13
13GPT-5 Mini OpenAI39.7%SWE-bench2026-02-13

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does SWE-bench Multilingual measure?

Real GitHub issues from repositories in nine programming languages other than Python, solved in the same bash-only agent.

Which model has the highest SWE-bench Multilingual score?

As of October 2026, Gemini 3 Flash Preview has the highest published SWE-bench Multilingual score on Noometry at 72.7%, out of 13 models with results.

What is the best open-weight model on SWE-bench Multilingual?

GLM-5 has the highest SWE-bench Multilingual accuracy among open-weight models at 69.7%, ranking 4 of 13 overall.