Coding benchmark

MBPP+ leaderboard

As of October 2026, o1 has the highest published MBPP+ score on Noometry at 80.2%, out of 38 models with results.

Last verified

About MBPP+

Entry-level Python problems from MBPP with extended test suites.

Category
Coding
Introduced
2023
Size
378 problems
Format
Code generation
Unit
Percent (random guessing ≈ 0%)
Official site
evalplus.github.io

Top 15 models

Top models on MBPP+
  1. o1 80.2%
  2. o1-mini 78.8%
  3. Qwen2.5-Coder-32B 77%
  4. Deepseek Coder v2 75.1%
  5. Gemini 1.5 Pro (May 2024) 74.6%
  6. Claude 3.5 Sonnet 74.3%
  7. DeepSeek-V2.5 (Sep 2024) 74.1%
  8. Claude 3 Opus 73.3%
  9. GPT-4 Turbo 73.3%
  10. DeepSeek-V3 73%
  11. GPT-4o 72.2%
  12. GPT-4o mini 72.2%
  13. DeepSeek Coder 33B 70.1%
  14. GPT-3.5-turbo 69.7%
  15. Claude 3 Sonnet 69.3%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

MBPP+ results by model
#ModelProviderScoreSettingSourceDate
1o1 OpenAI80.2%sept 2024EvalPlus
2o1-mini OpenAI78.8%sept 2024EvalPlus
3Qwen2.5-Coder-32B Alibaba (Qwen)77%EvalPlus
4Deepseek Coder v2 DeepSeek75.1%EvalPlus
5Gemini 1.5 Pro (May 2024) Google74.6%EvalPlus
6Claude 3.5 Sonnet Anthropic74.3%june 2024EvalPlus
7DeepSeek-V2.5 (Sep 2024) DeepSeek74.1%nov 2024EvalPlus
8Claude 3 Opus Anthropic73.3%mar 2024EvalPlus
9GPT-4 Turbo OpenAI73.3%nov 2023EvalPlus
10DeepSeek-V3 DeepSeek73%nov 2024EvalPlus
11GPT-4o OpenAI72.2%aug 2024EvalPlus
12GPT-4o mini OpenAI72.2%july 2024EvalPlus
13DeepSeek Coder 33B DeepSeek70.1%EvalPlus
14GPT-3.5-turbo OpenAI69.7%nov 2023EvalPlus
15Claude 3 Sonnet Anthropic69.3%mar 2024EvalPlus
16Llama 3-70B Meta69%EvalPlus
17Claude 3 Haiku Anthropic68.8%mar 2024EvalPlus
18Gemini 1.5 Flash (May 2024) Google67.5%EvalPlus
19DeepSeek Coder 6.7B DeepSeek65.6%EvalPlus
20Mixtral 8x22B Mistral AI64.3%EvalPlus
21Command R+ Cohere63.5%EvalPlus
22Codestral Mistral AI61.9%EvalPlus
23Qwen1.5-72B Alibaba (Qwen)61.6%EvalPlus
24Gemini 1.0 Pro Google61.4%EvalPlus
25Mistral Large Mistral AI59.5%mar 2024EvalPlus
26Codellama 34b Instruct Meta56.3%EvalPlus
27DBRX Databricks55.8%EvalPlus
28Llama 3.1-8B Meta55.6%EvalPlus
29DeepSeek Coder 1.3B DeepSeek54.8%EvalPlus
30Llama 3-8B Meta54.8%EvalPlus
31Phi-2 Microsoft54.2%EvalPlus
32Phi 3 Mini 4k Instruct Microsoft54.2%EvalPlus
33Mixtral 8x7B Mistral AI49.7%EvalPlus
34Gemma 1.1 7b IT Google45%EvalPlus
35Gemma 7B Google43.4%EvalPlus
36Mistral 7B Mistral AI42.1%EvalPlus
37Gemma 2B Google34.1%EvalPlus
38Gemma 1.1 2b IT Google23.3%EvalPlus

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does MBPP+ measure?

Entry-level Python problems from MBPP with extended test suites.

Which model has the highest MBPP+ score?

As of October 2026, o1 has the highest published MBPP+ score on Noometry at 80.2%, out of 38 models with results.

What is the best open-weight model on MBPP+?

Qwen2.5-Coder-32B has the highest MBPP+ accuracy among open-weight models at 77%, ranking 3 of 38 overall.