Coding benchmark
HumanEval+ leaderboard
As of October 2026, o1 has the highest published HumanEval+ score on Noometry at 89%, out of 45 models with results.
Last verified
About HumanEval+
HumanEval's 164 Python problems with many more test cases per problem, which catches solutions that only pass the original tests.
- Category
- Coding
- Introduced
- 2023
- Size
- 164 problems
- Format
- Code generation
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- evalplus.github.io
Top 15 models
- o1 89%
- o1-mini 89%
- GPT-4o 87.2%
- Qwen2.5-Coder-32B 87.2%
- DeepSeek-V3 86.6%
- GPT-4 Turbo 86.6%
- DeepSeek-V2.5 (Sep 2024) 83.5%
- GPT-4o mini 83.5%
- Deepseek Coder v2 82.3%
- Claude 3.5 Sonnet 81.7%
- Gemini 1.5 Pro (May 2024) 79.3%
- GPT-4 79.3%
- Claude 3 Opus 77.4%
- Gemini 1.5 Flash (May 2024) 75.6%
- DeepSeek Coder 33B 75%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does HumanEval+ measure?
HumanEval's 164 Python problems with many more test cases per problem, which catches solutions that only pass the original tests.
Which model has the highest HumanEval+ score?
As of October 2026, o1 has the highest published HumanEval+ score on Noometry at 89%, out of 45 models with results.
What is the best open-weight model on HumanEval+?
Qwen2.5-Coder-32B has the highest HumanEval+ accuracy among open-weight models at 87.2%, ranking 4 of 45 overall.