Reasoning benchmark
HellaSwag leaderboard
As of October 2026, GPT-4 has the highest published HellaSwag score on Noometry at 95.3%, out of 29 models with results.
Last verified
About HellaSwag
Commonsense sentence completion with adversarially filtered wrong endings.
- Category
- Reasoning
- Introduced
- 2019
- Format
- Multiple choice
- Unit
- Percent (random guessing ≈ 25%)
- Official site
- rowanzellers.com
Top 15 models
- GPT-4 95.3%
- GPT-4 95.3%
- Llama 3.1-405B 89.2%
- Falcon-180B 89%
- DeepSeek-V3 88.9%
- DeepSeek-V2 (MoE-236B, May 2024) 87.1%
- Mixtral 8x7B 86.7%
- Llama 2-70B 85.3%
- Falcon-40B 85.3%
- Qwen2.5 72B Instruct 84.8%
- Qwen2.5-Coder-32B 83%
- Falcon 2 11B 82.9%
- Nemotron-4 15B 82.4%
- phi-3-medium 14B 82.4%
- Gemma 7B 82.2%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does HellaSwag measure?
Commonsense sentence completion with adversarially filtered wrong endings.
Which model has the highest HellaSwag score?
As of October 2026, GPT-4 has the highest published HellaSwag score on Noometry at 95.3%, out of 29 models with results.
What is the best open-weight model on HellaSwag?
Llama 3.1-405B has the highest HellaSwag accuracy among open-weight models at 89.2%, ranking 3 of 29 overall.