Coding benchmark

BigCodeBench Complete leaderboard

As of October 2026, DeepSeek-V3 has the highest published BigCodeBench Complete score on Noometry at 62.2%, out of 66 models with results.

Last verified

About BigCodeBench Complete

The BigCodeBench tasks posed as docstring completion instead of instructions.

Category
Coding
Introduced
2024
Size
1,140 tasks
Format
Code completion
Unit
Percent (random guessing ≈ 0%)
Official site
bigcode-bench.github.io

Top 15 models

Top models on BigCodeBench Complete
  1. DeepSeek-V3 62.2%
  2. Llama 4 Maverick 61.4%
  3. GPT-4o 61.1%
  4. Gemini 2.0 Flash (Feb 2025) 59.9%
  5. Deepseek Coder v2 59.7%
  6. DeepSeek-V2 (MoE-236B, May 2024) 59.4%
  7. Claude 3.5 Haiku 59%
  8. Claude 3.5 Sonnet 58.6%
  9. GPT-4 Turbo 58.2%
  10. Qwen2.5-Coder-32B 58%
  11. Gemini 1.5 Pro (May 2024) 57.5%
  12. Llama-3.3-70B-Instruct 57.5%
  13. Claude 3 Opus 57.4%
  14. GPT-4o mini 57.4%
  15. GPT-4 57.2%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

BigCodeBench Complete results by model
#ModelProviderScoreSettingSourceDate
1DeepSeek-V3 DeepSeek62.2%BigCodeBench2024-12-26
2Llama 4 Maverick Meta61.4%BigCodeBench2025-04-05
3GPT-4o OpenAI61.1%BigCodeBench2024-05-13
4Gemini 2.0 Flash (Feb 2025) Google59.9%BigCodeBench2025-02-05
5Deepseek Coder v2 DeepSeek59.7%BigCodeBench2024-06-17
6DeepSeek-V2 (MoE-236B, May 2024) DeepSeek59.4%2024-06-28BigCodeBench2024-06-28
7Claude 3.5 Haiku Anthropic59%BigCodeBench2024-10-22
8Claude 3.5 Sonnet Anthropic58.6%BigCodeBench2024-06-20
9GPT-4 Turbo OpenAI58.2%BigCodeBench2024-04-09
10Qwen2.5-Coder-32B Alibaba (Qwen)58%BigCodeBench2024-09-19
11Gemini 1.5 Pro (May 2024) Google57.5%BigCodeBench2024-05-14
12Llama-3.3-70B-Instruct Meta57.5%BigCodeBench2024-12-19
13Claude 3 Opus Anthropic57.4%BigCodeBench2024-02-29
14GPT-4o mini OpenAI57.4%BigCodeBench2024-07-18
15GPT-4 OpenAI57.2%BigCodeBench2024-06-13
16Qwen2.5 72B Instruct Alibaba (Qwen)55.9%BigCodeBench2024-09-19
17Phi-4 Microsoft55.4%BigCodeBench2024-12-13
18Gemini 1.5 Flash (May 2024) Google55.1%BigCodeBench2024-05-14
19DeepSeek-R1-Distill-Qwen-32B DeepSeek54.9%BigCodeBench2025-01-20
20Llama 3.1-70B Meta54.8%BigCodeBench2024-07-23
21Llama 3-70B Meta54.5%BigCodeBench2024-04-18
22QwQ-32B Alibaba (Qwen)54.4%BigCodeBench2024-11-28
23Qwen2-72B Alibaba (Qwen)54%BigCodeBench2024-06-07
24Claude 3 Sonnet Anthropic53.8%BigCodeBench2024-02-29
25DeepSeek-V2.5 (Sep 2024) DeepSeek53.2%BigCodeBench2024-12-10
26Codestral Mistral AI52.5%BigCodeBench2024-05-23
27Gemma 2 27B Google52.5%BigCodeBench2024-06-19
28Qwen2.5 32B Instruct Alibaba (Qwen)52.3%BigCodeBench2024-09-19
29Qwen2.5 14B Instruct Alibaba (Qwen)52.2%BigCodeBench2024-09-19
30DeepSeek Coder 33B DeepSeek51.1%BigCodeBench2023-10-28
31GPT-3.5-turbo OpenAI50.6%BigCodeBench2024-01-25
32Mistral Small 3 Mistral AI50.4%BigCodeBench2025-01-31
33Mixtral 8x22B Mistral AI50.2%BigCodeBench2024-04-17
34Claude 3 Haiku Anthropic50.1%BigCodeBench2024-03-07
35DeepSeek-R1-Distill-Llama-70B DeepSeek49.9%BigCodeBench2025-01-20
36Codellama 70b Instruct Meta49.6%BigCodeBench2023-08-25
37phi-3-medium 14B Microsoft48.7%BigCodeBench2024-05-21
38DeepSeek-R1-Distill-Qwen-14B DeepSeek48.4%BigCodeBench2025-01-20
39Llama 3.1 Nemotron 70b Instruct NVIDIA48.2%BigCodeBench2024-09-25
40Yi-Large 01.AI47.2%BigCodeBench2024-05-13
41Mistral Small Mistral AI46.6%BigCodeBench2024-09-18
42Qwen2.5 7B Instruct Alibaba (Qwen)46.1%BigCodeBench2024-09-19
43Command R Cohere45.2%BigCodeBench2024-08-30
44Qwen1.5-110B Alibaba (Qwen)44.4%BigCodeBench2024-04-26
45DeepSeek Coder 6.7B DeepSeek43.8%BigCodeBench2023-10-28
46Yi-1.5-34B 01.AI43.8%BigCodeBench2024-05-20
47Llama 4 Scout Meta43.1%BigCodeBench2025-04-05
48Qwen1.5-32B Alibaba (Qwen)42%BigCodeBench2024-04-26
49Command R+ Cohere41.9%BigCodeBench2024-04-04
50Gemma 2 9B Google40.6%BigCodeBench2024-06-19
51Phi 3 Mini 128k Instruct Microsoft40.6%BigCodeBench2024-05-21
52Llama 3.1-8B Meta40.5%BigCodeBench2024-07-23
53Qwen1.5-72B Alibaba (Qwen)40.3%BigCodeBench2024-04-26
54Phi-3.5-mini Microsoft38.5%BigCodeBench2024-08-21
55StarCoder 2 15B NVIDIA38.4%BigCodeBench2024-02-29
56Mistral Large Mistral AI38.3%BigCodeBench2024-02-26
57Codellama 34b Instruct Meta37.1%BigCodeBench2023-08-25
58Llama 3-8B Meta36.9%BigCodeBench2024-04-18
59Granite 3.0 8b Instruct IBM35.4%BigCodeBench2024-10-21
60DeepSeek Coder 1.3B DeepSeek29.6%BigCodeBench2023-10-28
61Llama 3.2 3B Meta28.3%BigCodeBench2024-09-25
62StarCoder 2 7B NVIDIA27.7%BigCodeBench2024-02-29
63Mistral 7B Mistral AI27.3%BigCodeBench2024-05-22
64StarCoder 2 3B NVIDIA21.4%BigCodeBench2024-02-29
65Llama 3.2 1B Meta11.3%BigCodeBench2024-09-25
66DeepSeek-R1-Distill-Qwen-1.5B DeepSeek7.9%BigCodeBench2025-01-20

Compare the leaders

Other coding benchmarks

Frequently asked questions

What does BigCodeBench Complete measure?

The BigCodeBench tasks posed as docstring completion instead of instructions.

Which model has the highest BigCodeBench Complete score?

As of October 2026, DeepSeek-V3 has the highest published BigCodeBench Complete score on Noometry at 62.2%, out of 66 models with results.

What is the best open-weight model on BigCodeBench Complete?

DeepSeek-V3 has the highest BigCodeBench Complete accuracy among open-weight models at 62.2%, ranking 1 of 66 overall.