Coding benchmark

# HumanEval+ leaderboard

> HumanEval+ results for 45 AI models, led by o1 at 89%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/humaneval-plus
- Last updated: 2026-10-10
- Title: HumanEval+ Leaderboard (October 2026): Scores by Model

As of October 2026, o1 has the highest published HumanEval+ score on Noometry at 89%, out of 45 models with results.

Last verified October 10, 2026

## About HumanEval+

HumanEval's 164 Python problems with many more test cases per problem, which catches solutions that only pass the original tests.

- **Category:** [Coding](https://noometry.com/best/coding)
- **Introduced:** 2023
- **Size:** 164 problems
- **Format:** Code generation
- **Unit:** Percent (random guessing ≈ 0%)
- **Official site:** [evalplus.github.io](https://evalplus.github.io/leaderboard.html)

## Top 15 models

Top models on HumanEval+

1.  o1 89%
2.  o1-mini 89%
3.  GPT-4o 87.2%
4.  Qwen2.5-Coder-32B 87.2%
5.  DeepSeek-V3 86.6%
6.  GPT-4 Turbo 86.6%
7.  DeepSeek-V2.5 (Sep 2024) 83.5%
8.  GPT-4o mini 83.5%
9.  Deepseek Coder v2 82.3%
10.  Claude 3.5 Sonnet 81.7%
11.  Gemini 1.5 Pro (May 2024) 79.3%
12.  GPT-4 79.3%
13.  Claude 3 Opus 77.4%
14.  Gemini 1.5 Flash (May 2024) 75.6%
15.  DeepSeek Coder 33B 75%
16.  7075808590

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

HumanEval+ results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [o1](https://noometry.com/models/o1) | [OpenAI](https://noometry.com/providers/openai) | 89% | sept 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 2 | [o1-mini](https://noometry.com/models/o1-mini) | [OpenAI](https://noometry.com/providers/openai) | 89% | sept 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 3 | [GPT-4o](https://noometry.com/models/gpt-4o) | [OpenAI](https://noometry.com/providers/openai) | 87.2% | aug 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 4 | [Qwen2.5-Coder-32B](https://noometry.com/models/qwen2-5-coder-32b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 87.2% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 5 | [DeepSeek-V3](https://noometry.com/models/deepseek-v3) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 86.6% | nov 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 6 | [GPT-4 Turbo](https://noometry.com/models/gpt-4-turbo) | [OpenAI](https://noometry.com/providers/openai) | 86.6% | april 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 7 | [DeepSeek-V2.5 (Sep 2024)](https://noometry.com/models/deepseek-v2-5) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 83.5% | nov 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 8 | [GPT-4o mini](https://noometry.com/models/gpt-4o-mini) | [OpenAI](https://noometry.com/providers/openai) | 83.5% | july 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 9 | [Deepseek Coder v2](https://noometry.com/models/deepseek-coder-v2) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 82.3% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 10 | [Claude 3.5 Sonnet](https://noometry.com/models/claude-3-5-sonnet) | [Anthropic](https://noometry.com/providers/anthropic) | 81.7% | june 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 11 | [Gemini 1.5 Pro (May 2024)](https://noometry.com/models/gemini-1-5-pro) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 79.3% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 12 | [GPT-4](https://noometry.com/models/gpt-4) | [OpenAI](https://noometry.com/providers/openai) | 79.3% | may 2023 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 13 | [Claude 3 Opus](https://noometry.com/models/claude-3-opus) | [Anthropic](https://noometry.com/providers/anthropic) | 77.4% | mar 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 14 | [Gemini 1.5 Flash (May 2024)](https://noometry.com/models/gemini-1-5-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 75.6% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 15 | [DeepSeek Coder 33B](https://noometry.com/models/deepseek-coder-33b) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 75% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 16 | [Codestral](https://noometry.com/models/codestral) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 73.8% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 17 | [Llama 3-70B](https://noometry.com/models/llama-3-70b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 72% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 18 | [Mixtral 8x22B](https://noometry.com/models/mixtral-8x22b) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 72% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 19 | [DeepSeek Coder 6.7B](https://noometry.com/models/deepseek-coder-6-7b) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 71.3% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 20 | [GPT-3.5-turbo](https://noometry.com/models/gpt-3-5-turbo) | [OpenAI](https://noometry.com/providers/openai) | 70.7% | nov 2023 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 21 | [DBRX](https://noometry.com/models/dbrx) |  [![](/logos/databricks.svg) Databricks](https://noometry.com/providers/databricks) | 70.1% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 22 | [Claude 3 Haiku](https://noometry.com/models/claude-3-haiku) | [Anthropic](https://noometry.com/providers/anthropic) | 68.9% | mar 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 23 | [Codellama 70b Instruct](https://noometry.com/models/codellama-70b-instruct) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 65.9% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 24 | [Claude 3 Sonnet](https://noometry.com/models/claude-3-sonnet) | [Anthropic](https://noometry.com/providers/anthropic) | 64% | mar 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 25 | [Llama 3.1-8B](https://noometry.com/models/llama-3-1-8b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 62.8% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 26 | [Mistral Large](https://noometry.com/models/mistral-large) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 62.2% | mar 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 27 | [Claude 2](https://noometry.com/models/claude-2) | [Anthropic](https://noometry.com/providers/anthropic) | 61.6% | mar 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 28 | [DeepSeek Coder 1.3B](https://noometry.com/models/deepseek-coder-1-3b) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 60.4% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 29 | [Phi 3 Mini 4k Instruct](https://noometry.com/models/phi-3-mini-4k-instruct) |  [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft) | 59.1% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 30 | [Qwen1.5-72B](https://noometry.com/models/qwen1-5-72b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 59.1% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 31 | [Command R+](https://noometry.com/models/command-r-plus) |  [![](/logos/cohere.svg) Cohere](https://noometry.com/providers/cohere) | 56.7% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 32 | [Llama 3-8B](https://noometry.com/models/llama-3-8b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 56.7% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 33 | [Gemini 1.0 Pro](https://noometry.com/models/gemini-1-0-pro) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 55.5% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 34 | [Claude Instant](https://noometry.com/models/claude-instant) | [Anthropic](https://noometry.com/providers/anthropic) | 50.6% | mar 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 35 | [Phi-2](https://noometry.com/models/phi-2) |  [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft) | 45.1% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 36 | [Codellama 34b Instruct](https://noometry.com/models/codellama-34b-instruct) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 43.9% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 37 | [Mixtral 8x7B](https://noometry.com/models/mixtral-8x7b) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 39.6% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 38 | [StarCoder 2 15B](https://noometry.com/models/starcoder-2-15b) |  [![](/logos/nvidia.svg) NVIDIA](https://noometry.com/providers/nvidia) | 37.8% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 39 | [Mistral 7B](https://noometry.com/models/mistral-7b) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 36% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 40 | [Gemma 1.1 7b IT](https://noometry.com/models/gemma-1-1-7b-it) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 35.4% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 41 | [StarCoder 2 7B](https://noometry.com/models/starcoder-2-7b) |  [![](/logos/nvidia.svg) NVIDIA](https://noometry.com/providers/nvidia) | 29.9% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 42 | [Gemma 7B](https://noometry.com/models/gemma-7b) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 28.7% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 43 | [StarCoder 2 3B](https://noometry.com/models/starcoder-2-3b) |  [![](/logos/nvidia.svg) NVIDIA](https://noometry.com/providers/nvidia) | 27.4% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 44 | [Gemma 2B](https://noometry.com/models/gemma-2b) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 20.7% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| 45 | [Gemma 1.1 2b IT](https://noometry.com/models/gemma-1-1-2b-it) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 17.7% |  | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |

## Compare the leaders

-   [o1 vs o1-mini](https://noometry.com/compare/o1-vs-o1-mini)
-   [o1 vs GPT-4o](https://noometry.com/compare/gpt-4o-vs-o1)
-   [o1 vs Qwen2.5-Coder-32B](https://noometry.com/compare/o1-vs-qwen2-5-coder-32b)
-   [o1 vs DeepSeek-V3](https://noometry.com/compare/deepseek-v3-vs-o1)
-   [o1-mini vs GPT-4o](https://noometry.com/compare/gpt-4o-vs-o1-mini)
-   [o1-mini vs Qwen2.5-Coder-32B](https://noometry.com/compare/o1-mini-vs-qwen2-5-coder-32b)

## Other coding benchmarks

-   [SWE-bench Verified](https://noometry.com/benchmarks/swe-bench-verified)
-   [DeepSWE](https://noometry.com/benchmarks/deepswe)
-   [FrontierCode](https://noometry.com/benchmarks/frontiercode)
-   [SWE-bench Verified (bash only)](https://noometry.com/benchmarks/swe-bench-bash-only)
-   [Aider Polyglot](https://noometry.com/benchmarks/aider-polyglot)
-   [LMArena WebDev](https://noometry.com/benchmarks/arena-webdev)
-   [CursorBench](https://noometry.com/benchmarks/cursorbench)
-   [SWE-bench Multilingual](https://noometry.com/benchmarks/swe-bench-multilingual)
-   [FrontierSWE](https://noometry.com/benchmarks/frontierswe)
-   [SciCode](https://noometry.com/benchmarks/scicode)
-   [GSO](https://noometry.com/benchmarks/gso-bench)
-   [WeirdML](https://noometry.com/benchmarks/weirdml)

## Frequently asked questions

### What does HumanEval+ measure?

HumanEval's 164 Python problems with many more test cases per problem, which catches solutions that only pass the original tests.

### Which model has the highest HumanEval+ score?

As of October 2026, o1 has the highest published HumanEval+ score on Noometry at 89%, out of 45 models with results.

### What is the best open-weight model on HumanEval+?

Qwen2.5-Coder-32B has the highest HumanEval+ accuracy among open-weight models at 87.2%, ranking 4 of 45 overall.

### Cite this page

Noometry. (2026). HumanEval+ leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/humaneval-plus

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/humaneval-plus.md).
