Reasoning benchmark
LiveBench Reasoning leaderboard
As of October 2026, GPT-5.1 has the highest published LiveBench Reasoning score on Noometry at 95.8%, out of 39 models with results.
Last verified
About LiveBench Reasoning
LiveBench reasoning tasks such as logic puzzles, refreshed to limit contamination.
- Category
- Reasoning
- Introduced
- 2024
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- livebench.ai
Top 15 models
- GPT-5.1 95.8%
- o1 91.6%
- Gemini 2.5 Pro 89.8%
- o3-mini 89.6%
- Claude 3.7 Sonnet 87.8%
- QwQ-32B 83.5%
- DeepSeek-R1 83.2%
- Gemini 2.0 Flash (Feb 2025) 78.2%
- o1-mini 72.3%
- GPT-4.5 71.1%
- DeepSeek-R1-Distill-Llama-70B 67.6%
- DeepSeek-V3 65.8%
- Gemini 2.0 Pro 60.1%
- Claude 3.5 Sonnet 56.7%
- GPT-4o 55.8%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does LiveBench Reasoning measure?
LiveBench reasoning tasks such as logic puzzles, refreshed to limit contamination.
Which model has the highest LiveBench Reasoning score?
As of October 2026, GPT-5.1 has the highest published LiveBench Reasoning score on Noometry at 95.8%, out of 39 models with results.
What is the best open-weight model on LiveBench Reasoning?
QwQ-32B has the highest LiveBench Reasoning accuracy among open-weight models at 83.5%, ranking 6 of 39 overall.