Reasoning benchmark
WinoGrande leaderboard
As of October 2026, Llama 3.1-405B has the highest published WinoGrande score on Noometry at 89.2%, out of 43 models with results.
Last verified
About WinoGrande
Pronoun-resolution problems that need commonsense reasoning.
- Category
- Reasoning
- Introduced
- 2019
- Format
- Binary choice
- Unit
- Percent (random guessing ≈ 50%)
- Official site
- winogrande.allenai.org
Top 15 models
- Llama 3.1-405B 89.2%
- Claude 3 Opus 88.5%
- GPT-4 87.5%
- GPT-4 87.5%
- Falcon-180B 87.1%
- DeepSeek-V2 (MoE-236B, May 2024) 86.3%
- DeepSeek-V3 85.2%
- Deepseek Coder v2 83.7%
- Llama 3-70B 83.5%
- Qwen2.5 72B Instruct 82.3%
- GPT-3.5-turbo 81.6%
- phi-3-medium 14B 81.5%
- Phi 3 Small 8k Instruct 81.5%
- Qwen2.5-Coder-32B 80.8%
- Llama 2-70B 80.2%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does WinoGrande measure?
Pronoun-resolution problems that need commonsense reasoning.
Which model has the highest WinoGrande score?
As of October 2026, Llama 3.1-405B has the highest published WinoGrande score on Noometry at 89.2%, out of 43 models with results.
What is the best open-weight model on WinoGrande?
Llama 3.1-405B has the highest WinoGrande accuracy among open-weight models at 89.2%, ranking 1 of 43 overall.