Coding benchmark
BigCodeBench Complete leaderboard
As of October 2026, DeepSeek-V3 has the highest published BigCodeBench Complete score on Noometry at 62.2%, out of 66 models with results.
Last verified
About BigCodeBench Complete
The BigCodeBench tasks posed as docstring completion instead of instructions.
- Category
- Coding
- Introduced
- 2024
- Size
- 1,140 tasks
- Format
- Code completion
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- bigcode-bench.github.io
Top 15 models
- DeepSeek-V3 62.2%
- Llama 4 Maverick 61.4%
- GPT-4o 61.1%
- Gemini 2.0 Flash (Feb 2025) 59.9%
- Deepseek Coder v2 59.7%
- DeepSeek-V2 (MoE-236B, May 2024) 59.4%
- Claude 3.5 Haiku 59%
- Claude 3.5 Sonnet 58.6%
- GPT-4 Turbo 58.2%
- Qwen2.5-Coder-32B 58%
- Gemini 1.5 Pro (May 2024) 57.5%
- Llama-3.3-70B-Instruct 57.5%
- Claude 3 Opus 57.4%
- GPT-4o mini 57.4%
- GPT-4 57.2%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does BigCodeBench Complete measure?
The BigCodeBench tasks posed as docstring completion instead of instructions.
Which model has the highest BigCodeBench Complete score?
As of October 2026, DeepSeek-V3 has the highest published BigCodeBench Complete score on Noometry at 62.2%, out of 66 models with results.
What is the best open-weight model on BigCodeBench Complete?
DeepSeek-V3 has the highest BigCodeBench Complete accuracy among open-weight models at 62.2%, ranking 1 of 66 overall.