Coding benchmark
MBPP+ leaderboard
As of October 2026, o1 has the highest published MBPP+ score on Noometry at 80.2%, out of 38 models with results.
Last verified
About MBPP+
Entry-level Python problems from MBPP with extended test suites.
- Category
- Coding
- Introduced
- 2023
- Size
- 378 problems
- Format
- Code generation
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- evalplus.github.io
Top 15 models
- o1 80.2%
- o1-mini 78.8%
- Qwen2.5-Coder-32B 77%
- Deepseek Coder v2 75.1%
- Gemini 1.5 Pro (May 2024) 74.6%
- Claude 3.5 Sonnet 74.3%
- DeepSeek-V2.5 (Sep 2024) 74.1%
- Claude 3 Opus 73.3%
- GPT-4 Turbo 73.3%
- DeepSeek-V3 73%
- GPT-4o 72.2%
- GPT-4o mini 72.2%
- DeepSeek Coder 33B 70.1%
- GPT-3.5-turbo 69.7%
- Claude 3 Sonnet 69.3%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does MBPP+ measure?
Entry-level Python problems from MBPP with extended test suites.
Which model has the highest MBPP+ score?
As of October 2026, o1 has the highest published MBPP+ score on Noometry at 80.2%, out of 38 models with results.
What is the best open-weight model on MBPP+?
Qwen2.5-Coder-32B has the highest MBPP+ accuracy among open-weight models at 77%, ranking 3 of 38 overall.