Coding benchmark

# BigCodeBench Instruct leaderboard

> BigCodeBench Instruct results for 64 AI models, led by GPT-4o at 51.1%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/bigcodebench-instruct
- Last updated: 2026-10-10
- Title: BigCodeBench Instruct Leaderboard (October 2026): Scores by Model

As of October 2026, GPT-4o has the highest published BigCodeBench Instruct score on Noometry at 51.1%, out of 64 models with results.

Last verified October 10, 2026

## About BigCodeBench Instruct

Practical Python tasks that call functions from 139 libraries, written from natural-language instructions.

- **Category:** [Coding](https://noometry.com/best/coding)
- **Introduced:** 2024
- **Size:** 1,140 tasks
- **Format:** Code generation
- **Unit:** Percent (random guessing ≈ 0%)
- **Official site:** [bigcode-bench.github.io](https://bigcode-bench.github.io/)

## Top 15 models

Top models on BigCodeBench Instruct

1.  GPT-4o 51.1%
2.  DeepSeek-V3 50%
3.  Llama 4 Maverick 49.7%
4.  Qwen2.5-Coder-32B 49%
5.  DeepSeek-V2 (MoE-236B, May 2024) 48.9%
6.  GPT-4.1 mini 48.9%
7.  DeepSeek-V2.5 (Sep 2024) 48.6%
8.  Deepseek Coder v2 48.2%
9.  GPT-4 Turbo 48.2%
10.  Llama-3.3-70B-Instruct 46.9%
11.  Claude 3.5 Sonnet 46.8%
12.  Claude 3.5 Haiku 46.1%
13.  GPT-4o mini 46.1%
14.  Llama 3.1-70B 46.1%
15.  GPT-4 46%
16.  4446485052

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

BigCodeBench Instruct results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GPT-4o](https://noometry.com/models/gpt-4o) | [OpenAI](https://noometry.com/providers/openai) | 51.1% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-05-13 |
| 2 | [DeepSeek-V3](https://noometry.com/models/deepseek-v3) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 50% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-12-26 |
| 3 | [Llama 4 Maverick](https://noometry.com/models/llama-4-maverick) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 49.7% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2025-04-05 |
| 4 | [Qwen2.5-Coder-32B](https://noometry.com/models/qwen2-5-coder-32b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 49% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-19 |
| 5 | [DeepSeek-V2 (MoE-236B, May 2024)](https://noometry.com/models/deepseek-v2) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 48.9% | 2024-06-28 | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-06-28 |
| 6 | [GPT-4.1 mini](https://noometry.com/models/gpt-4-1-mini) | [OpenAI](https://noometry.com/providers/openai) | 48.9% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2025-04-14 |
| 7 | [DeepSeek-V2.5 (Sep 2024)](https://noometry.com/models/deepseek-v2-5) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 48.6% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-12-10 |
| 8 | [Deepseek Coder v2](https://noometry.com/models/deepseek-coder-v2) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 48.2% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-06-17 |
| 9 | [GPT-4 Turbo](https://noometry.com/models/gpt-4-turbo) | [OpenAI](https://noometry.com/providers/openai) | 48.2% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-04-09 |
| 10 | [Llama-3.3-70B-Instruct](https://noometry.com/models/llama-3-3-70b-instruct) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 46.9% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-12-19 |
| 11 | [Claude 3.5 Sonnet](https://noometry.com/models/claude-3-5-sonnet) | [Anthropic](https://noometry.com/providers/anthropic) | 46.8% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-06-20 |
| 12 | [Claude 3.5 Haiku](https://noometry.com/models/claude-3-5-haiku) | [Anthropic](https://noometry.com/providers/anthropic) | 46.1% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-10-22 |
| 13 | [GPT-4o mini](https://noometry.com/models/gpt-4o-mini) | [OpenAI](https://noometry.com/providers/openai) | 46.1% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-07-18 |
| 14 | [Llama 3.1-70B](https://noometry.com/models/llama-3-1-70b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 46.1% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-07-23 |
| 15 | [GPT-4](https://noometry.com/models/gpt-4) | [OpenAI](https://noometry.com/providers/openai) | 46% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-06-13 |
| 16 | [Gemini 2.0 Flash (Feb 2025)](https://noometry.com/models/gemini-2-0-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 45.9% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2025-02-05 |
| 17 | [Qwen2.5 72B Instruct](https://noometry.com/models/qwen2-5-72b-instruct) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 45.8% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-19 |
| 18 | [Claude 3 Opus](https://noometry.com/models/claude-3-opus) | [Anthropic](https://noometry.com/providers/anthropic) | 45.5% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-02-29 |
| 19 | [Phi-4](https://noometry.com/models/phi-4) |  [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft) | 45.5% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-12-13 |
| 20 | [Mistral Small 3](https://noometry.com/models/mistral-small-3) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 45.3% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2025-01-31 |
| 21 | [Qwen2.5 32B Instruct](https://noometry.com/models/qwen2-5-32b-instruct) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 45% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-19 |
| 22 | [QwQ-32B](https://noometry.com/models/qwq-32b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 44.6% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-11-28 |
| 23 | [DeepSeek-R1-Distill-Qwen-32B](https://noometry.com/models/deepseek-r1-distill-qwen-32b) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 43.9% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2025-01-20 |
| 24 | [Gemini 1.5 Pro (May 2024)](https://noometry.com/models/gemini-1-5-pro) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 43.8% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-05-14 |
| 25 | [Llama 3-70B](https://noometry.com/models/llama-3-70b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 43.6% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-04-18 |
| 26 | [Gemini 1.5 Flash (May 2024)](https://noometry.com/models/gemini-1-5-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 43.5% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-05-14 |
| 27 | [Gemma 2 27B](https://noometry.com/models/gemma-2-27b) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 42.8% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-06-19 |
| 28 | [Claude 3 Sonnet](https://noometry.com/models/claude-3-sonnet) | [Anthropic](https://noometry.com/providers/anthropic) | 42.7% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-02-29 |
| 29 | [DeepSeek Coder 33B](https://noometry.com/models/deepseek-coder-33b) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 42% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2023-10-28 |
| 30 | [Codestral](https://noometry.com/models/codestral) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 41.8% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-05-23 |
| 31 | [Codellama 70b Instruct](https://noometry.com/models/codellama-70b-instruct) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 40.7% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2023-08-25 |
| 32 | [Mixtral 8x22B](https://noometry.com/models/mixtral-8x22b) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 40.6% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-04-17 |
| 33 | [Qwen2.5 14B Instruct](https://noometry.com/models/qwen2-5-14b-instruct) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 39.8% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-19 |
| 34 | [Claude 3 Haiku](https://noometry.com/models/claude-3-haiku) | [Anthropic](https://noometry.com/providers/anthropic) | 39.4% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-03-07 |
| 35 | [GPT-3.5-turbo](https://noometry.com/models/gpt-3-5-turbo) | [OpenAI](https://noometry.com/providers/openai) | 39.1% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-01-25 |
| 36 | [Llama 3.1 Nemotron 70b Instruct](https://noometry.com/models/llama-3-1-nemotron-70b-instruct) |  [![](/logos/nvidia.svg) NVIDIA](https://noometry.com/providers/nvidia) | 38.7% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-25 |
| 37 | [Qwen2-72B](https://noometry.com/models/qwen2-72b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 38.5% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-06-07 |
| 38 | [DeepSeek-R1-Distill-Qwen-14B](https://noometry.com/models/deepseek-r1-distill-qwen-14b) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 38.1% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2025-01-20 |
| 39 | [Yi-Large](https://noometry.com/models/yi-large) | [01.AI](https://noometry.com/providers/01-ai) | 37.7% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-05-13 |
| 40 | [phi-3-medium 14B](https://noometry.com/models/phi-3-medium-14b) |  [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft) | 37.6% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-05-21 |
| 41 | [Qwen2.5 7B Instruct](https://noometry.com/models/qwen2-5-7b-instruct) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 37.6% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-19 |
| 42 | [Command R](https://noometry.com/models/command-r) |  [![](/logos/cohere.svg) Cohere](https://noometry.com/providers/cohere) | 37.1% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-08-30 |
| 43 | [Mistral Small](https://noometry.com/models/mistral-small) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 36.1% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-18 |
| 44 | [DeepSeek Coder 6.7B](https://noometry.com/models/deepseek-coder-6-7b) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 35.5% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2023-10-28 |
| 45 | [DeepSeek-R1-Distill-Llama-70B](https://noometry.com/models/deepseek-r1-distill-llama-70b) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 35.3% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2025-01-20 |
| 46 | [Qwen1.5-110B](https://noometry.com/models/qwen1-5-110b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 35% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-04-26 |
| 47 | [Gemma 2 9B](https://noometry.com/models/gemma-2-9b) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 34.7% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-06-19 |
| 48 | [Yi-1.5-34B](https://noometry.com/models/yi-1-5-34b) | [01.AI](https://noometry.com/providers/01-ai) | 33.9% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-05-20 |
| 49 | [Command R+](https://noometry.com/models/command-r-plus) |  [![](/logos/cohere.svg) Cohere](https://noometry.com/providers/cohere) | 33.8% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-04-04 |
| 50 | [Qwen1.5-72B](https://noometry.com/models/qwen1-5-72b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 33.2% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-04-26 |
| 51 | [Llama 3.1-8B](https://noometry.com/models/llama-3-1-8b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 32.8% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-07-23 |
| 52 | [Phi-3.5-mini](https://noometry.com/models/phi-3-5-mini) |  [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft) | 32.8% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-08-21 |
| 53 | [Qwen1.5-32B](https://noometry.com/models/qwen1-5-32b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 32.3% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-04-26 |
| 54 | [Llama 3-8B](https://noometry.com/models/llama-3-8b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 31.9% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-04-18 |
| 55 | [Mistral Large](https://noometry.com/models/mistral-large) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 30% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-02-26 |
| 56 | [Phi 3 Mini 128k Instruct](https://noometry.com/models/phi-3-mini-128k-instruct) |  [![](/logos/microsoft.svg) Microsoft](https://noometry.com/providers/microsoft) | 29.6% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-05-21 |
| 57 | [Granite 3.0 8b Instruct](https://noometry.com/models/granite-3-0-8b-instruct) | [IBM](https://noometry.com/providers/ibm) | 29.3% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-10-21 |
| 58 | [Codellama 34b Instruct](https://noometry.com/models/codellama-34b-instruct) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 29% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2023-08-25 |
| 59 | [Llama 3.2 3B](https://noometry.com/models/llama-3-2-3b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 23.4% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-25 |
| 60 | [DeepSeek Coder 1.3B](https://noometry.com/models/deepseek-coder-1-3b) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 22.8% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2023-10-28 |
| 61 | [Granite 3.0 2b Instruct](https://noometry.com/models/granite-3-0-2b-instruct) | [IBM](https://noometry.com/providers/ibm) | 20.5% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-10-21 |
| 62 | [Mistral 7B](https://noometry.com/models/mistral-7b) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 19.5% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-05-22 |
| 63 | [Llama 3.2 1B](https://noometry.com/models/llama-3-2-1b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 8.2% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-25 |
| 64 | [DeepSeek-R1-Distill-Qwen-1.5B](https://noometry.com/models/deepseek-r1-distill-qwen-1-5b) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 7% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2025-01-20 |

## Compare the leaders

-   [GPT-4o vs DeepSeek-V3](https://noometry.com/compare/deepseek-v3-vs-gpt-4o)
-   [GPT-4o vs Llama 4 Maverick](https://noometry.com/compare/gpt-4o-vs-llama-4-maverick)
-   [GPT-4o vs Qwen2.5-Coder-32B](https://noometry.com/compare/gpt-4o-vs-qwen2-5-coder-32b)
-   [GPT-4o vs DeepSeek-V2 (MoE-236B, May 2024)](https://noometry.com/compare/deepseek-v2-vs-gpt-4o)
-   [DeepSeek-V3 vs Llama 4 Maverick](https://noometry.com/compare/deepseek-v3-vs-llama-4-maverick)
-   [DeepSeek-V3 vs Qwen2.5-Coder-32B](https://noometry.com/compare/deepseek-v3-vs-qwen2-5-coder-32b)

## Other coding benchmarks

-   [SWE-bench Verified](https://noometry.com/benchmarks/swe-bench-verified)
-   [DeepSWE](https://noometry.com/benchmarks/deepswe)
-   [FrontierCode](https://noometry.com/benchmarks/frontiercode)
-   [SWE-bench Verified (bash only)](https://noometry.com/benchmarks/swe-bench-bash-only)
-   [Aider Polyglot](https://noometry.com/benchmarks/aider-polyglot)
-   [LMArena WebDev](https://noometry.com/benchmarks/arena-webdev)
-   [CursorBench](https://noometry.com/benchmarks/cursorbench)
-   [SWE-bench Multilingual](https://noometry.com/benchmarks/swe-bench-multilingual)
-   [FrontierSWE](https://noometry.com/benchmarks/frontierswe)
-   [SciCode](https://noometry.com/benchmarks/scicode)
-   [GSO](https://noometry.com/benchmarks/gso-bench)
-   [WeirdML](https://noometry.com/benchmarks/weirdml)

## Frequently asked questions

### What does BigCodeBench Instruct measure?

Practical Python tasks that call functions from 139 libraries, written from natural-language instructions.

### Which model has the highest BigCodeBench Instruct score?

As of October 2026, GPT-4o has the highest published BigCodeBench Instruct score on Noometry at 51.1%, out of 64 models with results.

### What is the best open-weight model on BigCodeBench Instruct?

DeepSeek-V3 has the highest BigCodeBench Instruct accuracy among open-weight models at 50%, ranking 2 of 64 overall.

### Cite this page

Noometry. (2026). BigCodeBench Instruct leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/bigcodebench-instruct

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/bigcodebench-instruct.md).
