Coding benchmark

BigCodeBench Instruct leaderboard

As of October 2026, GPT-4o has the highest published BigCodeBench Instruct score on Noometry at 51.1%, out of 64 models with results.

Last verified

About BigCodeBench Instruct

Practical Python tasks that call functions from 139 libraries, written from natural-language instructions.

Category
Coding
Introduced
2024
Size
1,140 tasks
Format
Code generation
Unit
Percent (random guessing ≈ 0%)
Official site
bigcode-bench.github.io

Top 15 models

Top models on BigCodeBench Instruct
  1. GPT-4o 51.1%
  2. DeepSeek-V3 50%
  3. Llama 4 Maverick 49.7%
  4. Qwen2.5-Coder-32B 49%
  5. DeepSeek-V2 (MoE-236B, May 2024) 48.9%
  6. GPT-4.1 mini 48.9%
  7. DeepSeek-V2.5 (Sep 2024) 48.6%
  8. Deepseek Coder v2 48.2%
  9. GPT-4 Turbo 48.2%
  10. Llama-3.3-70B-Instruct 46.9%
  11. Claude 3.5 Sonnet 46.8%
  12. Claude 3.5 Haiku 46.1%
  13. GPT-4o mini 46.1%
  14. Llama 3.1-70B 46.1%
  15. GPT-4 46%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

BigCodeBench Instruct results by model
#ModelProviderScoreSettingSourceDate
1GPT-4o OpenAI51.1%BigCodeBench2024-05-13
2DeepSeek-V3 DeepSeek50%BigCodeBench2024-12-26
3Llama 4 Maverick Meta49.7%BigCodeBench2025-04-05
4Qwen2.5-Coder-32B Alibaba (Qwen)49%BigCodeBench2024-09-19
5DeepSeek-V2 (MoE-236B, May 2024) DeepSeek48.9%2024-06-28BigCodeBench2024-06-28
6GPT-4.1 mini OpenAI48.9%BigCodeBench2025-04-14
7DeepSeek-V2.5 (Sep 2024) DeepSeek48.6%BigCodeBench2024-12-10
8Deepseek Coder v2 DeepSeek48.2%BigCodeBench2024-06-17
9GPT-4 Turbo OpenAI48.2%BigCodeBench2024-04-09
10Llama-3.3-70B-Instruct Meta46.9%BigCodeBench2024-12-19
11Claude 3.5 Sonnet Anthropic46.8%BigCodeBench2024-06-20
12Claude 3.5 Haiku Anthropic46.1%BigCodeBench2024-10-22
13GPT-4o mini OpenAI46.1%BigCodeBench2024-07-18
14Llama 3.1-70B Meta46.1%BigCodeBench2024-07-23
15GPT-4 OpenAI46%BigCodeBench2024-06-13
16Gemini 2.0 Flash (Feb 2025) Google45.9%BigCodeBench2025-02-05
17Qwen2.5 72B Instruct Alibaba (Qwen)45.8%BigCodeBench2024-09-19
18Claude 3 Opus Anthropic45.5%BigCodeBench2024-02-29
19Phi-4 Microsoft45.5%BigCodeBench2024-12-13
20Mistral Small 3 Mistral AI45.3%BigCodeBench2025-01-31
21Qwen2.5 32B Instruct Alibaba (Qwen)45%BigCodeBench2024-09-19
22QwQ-32B Alibaba (Qwen)44.6%BigCodeBench2024-11-28
23DeepSeek-R1-Distill-Qwen-32B DeepSeek43.9%BigCodeBench2025-01-20
24Gemini 1.5 Pro (May 2024) Google43.8%BigCodeBench2024-05-14
25Llama 3-70B Meta43.6%BigCodeBench2024-04-18
26Gemini 1.5 Flash (May 2024) Google43.5%BigCodeBench2024-05-14
27Gemma 2 27B Google42.8%BigCodeBench2024-06-19
28Claude 3 Sonnet Anthropic42.7%BigCodeBench2024-02-29
29DeepSeek Coder 33B DeepSeek42%BigCodeBench2023-10-28
30Codestral Mistral AI41.8%BigCodeBench2024-05-23
31Codellama 70b Instruct Meta40.7%BigCodeBench2023-08-25
32Mixtral 8x22B Mistral AI40.6%BigCodeBench2024-04-17
33Qwen2.5 14B Instruct Alibaba (Qwen)39.8%BigCodeBench2024-09-19
34Claude 3 Haiku Anthropic39.4%BigCodeBench2024-03-07
35GPT-3.5-turbo OpenAI39.1%BigCodeBench2024-01-25
36Llama 3.1 Nemotron 70b Instruct NVIDIA38.7%BigCodeBench2024-09-25
37Qwen2-72B Alibaba (Qwen)38.5%BigCodeBench2024-06-07
38DeepSeek-R1-Distill-Qwen-14B DeepSeek38.1%BigCodeBench2025-01-20
39Yi-Large 01.AI37.7%BigCodeBench2024-05-13
40phi-3-medium 14B Microsoft37.6%BigCodeBench2024-05-21
41Qwen2.5 7B Instruct Alibaba (Qwen)37.6%BigCodeBench2024-09-19
42Command R Cohere37.1%BigCodeBench2024-08-30
43Mistral Small Mistral AI36.1%BigCodeBench2024-09-18
44DeepSeek Coder 6.7B DeepSeek35.5%BigCodeBench2023-10-28
45DeepSeek-R1-Distill-Llama-70B DeepSeek35.3%BigCodeBench2025-01-20
46Qwen1.5-110B Alibaba (Qwen)35%BigCodeBench2024-04-26
47Gemma 2 9B Google34.7%BigCodeBench2024-06-19
48Yi-1.5-34B 01.AI33.9%BigCodeBench2024-05-20
49Command R+ Cohere33.8%BigCodeBench2024-04-04
50Qwen1.5-72B Alibaba (Qwen)33.2%BigCodeBench2024-04-26
51Llama 3.1-8B Meta32.8%BigCodeBench2024-07-23
52Phi-3.5-mini Microsoft32.8%BigCodeBench2024-08-21
53Qwen1.5-32B Alibaba (Qwen)32.3%BigCodeBench2024-04-26
54Llama 3-8B Meta31.9%BigCodeBench2024-04-18
55Mistral Large Mistral AI30%BigCodeBench2024-02-26
56Phi 3 Mini 128k Instruct Microsoft29.6%BigCodeBench2024-05-21
57Granite 3.0 8b Instruct IBM29.3%BigCodeBench2024-10-21
58Codellama 34b Instruct Meta29%BigCodeBench2023-08-25
59Llama 3.2 3B Meta23.4%BigCodeBench2024-09-25
60DeepSeek Coder 1.3B DeepSeek22.8%BigCodeBench2023-10-28
61Granite 3.0 2b Instruct IBM20.5%BigCodeBench2024-10-21
62Mistral 7B Mistral AI19.5%BigCodeBench2024-05-22
63Llama 3.2 1B Meta8.2%BigCodeBench2024-09-25
64DeepSeek-R1-Distill-Qwen-1.5B DeepSeek7%BigCodeBench2025-01-20

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does BigCodeBench Instruct measure?

Practical Python tasks that call functions from 139 libraries, written from natural-language instructions.

Which model has the highest BigCodeBench Instruct score?

As of October 2026, GPT-4o has the highest published BigCodeBench Instruct score on Noometry at 51.1%, out of 64 models with results.

What is the best open-weight model on BigCodeBench Instruct?

DeepSeek-V3 has the highest BigCodeBench Instruct accuracy among open-weight models at 50%, ranking 2 of 64 overall.