Reasoning benchmark

WinoGrande leaderboard

As of October 2026, Llama 3.1-405B has the highest published WinoGrande score on Noometry at 89.2%, out of 43 models with results.

Last verified

About WinoGrande

Pronoun-resolution problems that need commonsense reasoning.

Category
Reasoning
Introduced
2019
Format
Binary choice
Unit
Percent (random guessing ≈ 50%)
Official site
winogrande.allenai.org

Top 15 models

Top models on WinoGrande
  1. Llama 3.1-405B 89.2%
  2. Claude 3 Opus 88.5%
  3. GPT-4 87.5%
  4. GPT-4 87.5%
  5. Falcon-180B 87.1%
  6. DeepSeek-V2 (MoE-236B, May 2024) 86.3%
  7. DeepSeek-V3 85.2%
  8. Deepseek Coder v2 83.7%
  9. Llama 3-70B 83.5%
  10. Qwen2.5 72B Instruct 82.3%
  11. GPT-3.5-turbo 81.6%
  12. phi-3-medium 14B 81.5%
  13. Phi 3 Small 8k Instruct 81.5%
  14. Qwen2.5-Coder-32B 80.8%
  15. Llama 2-70B 80.2%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

WinoGrande results by model
#ModelProviderScoreSettingSourceDate
1Llama 3.1-405B Meta89.2%Epoch AI
2Claude 3 Opus Anthropic88.5%Epoch AI
3GPT-4 OpenAI87.5%Epoch AI
4GPT-4 OpenAI87.5%Epoch AI
5Falcon-180B Technology Innovation Institute87.1%Epoch AI
6DeepSeek-V2 (MoE-236B, May 2024) DeepSeek86.3%Epoch AI
7DeepSeek-V3 DeepSeek85.2%Epoch AI
8Deepseek Coder v2 DeepSeek83.7%Epoch AI
9Llama 3-70B Meta83.5%Epoch AI
10Qwen2.5 72B Instruct Alibaba (Qwen)82.3%Epoch AI
11GPT-3.5-turbo OpenAI81.6%Epoch AI
12phi-3-medium 14B Microsoft81.5%Epoch AI
13Phi 3 Small 8k Instruct Microsoft81.5%Epoch AI
14Qwen2.5-Coder-32B Alibaba (Qwen)80.8%Epoch AI
15Llama 2-70B Meta80.2%Epoch AI
16Gemma 7B Google79%Epoch AI
17Falcon 2 11B Technology Innovation Institute78.3%Epoch AI
18Nemotron-4 15B NVIDIA78%Epoch AI
19Mixtral 8x7B Mistral AI77.2%Epoch AI
20Falcon-40B Technology Innovation Institute76.9%Epoch AI
21Llama 2-34B Meta76.7%Epoch AI
22Llama 3-8B Meta75.7%Epoch AI
23Mistral 7B Mistral AI75.3%Epoch AI
24Claude 3 Sonnet Anthropic75.1%Epoch AI
25Claude 3 Haiku Anthropic74.2%Epoch AI
26Phi-1.5 Microsoft73.4%5Epoch AI
27Llama 13b Meta73%Epoch AI
28Qwen2.5-Coder (1.5B) Alibaba (Qwen)72.9%Epoch AI
29Llama 2-13B Meta72.8%Epoch AI
30Yi 6B 01.AI71.3%Epoch AI
31Phi 3 Mini 4k Instruct Microsoft70.8%Epoch AI
32Llama 2-7B Meta69.2%Epoch AI
33Falcon-7B Technology Innovation Institute67.2%Epoch AI
34INTELLECT-1 Hugging Face65.8%Epoch AI
35Gemma 2B Google65.4%Epoch AI
36StarCoder 2 15B NVIDIA64.3%Epoch AI
37DeepSeek Coder 33B DeepSeek62%Epoch AI
38Dolly 2.0-12b Databricks61.8%Epoch AI
39DeepSeek Coder 6.7B DeepSeek57.6%Epoch AI
40StarCoder 2 3B NVIDIA57.1%Epoch AI
41StarCoder 2 7B NVIDIA57.1%Epoch AI
42Phi-2 Microsoft54.7%Epoch AI
43DeepSeek Coder 1.3B DeepSeek53.3%Epoch AI

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does WinoGrande measure?

Pronoun-resolution problems that need commonsense reasoning.

Which model has the highest WinoGrande score?

As of October 2026, Llama 3.1-405B has the highest published WinoGrande score on Noometry at 89.2%, out of 43 models with results.

What is the best open-weight model on WinoGrande?

Llama 3.1-405B has the highest WinoGrande accuracy among open-weight models at 89.2%, ranking 1 of 43 overall.