Writing & Preference benchmark
LiveBench Language leaderboard
As of October 2026, GPT-5.1 has the highest published LiveBench Language score on Noometry at 80.2%, out of 39 models with results.
Last verified
About LiveBench Language
LiveBench language tasks such as typo fixing and plot unscrambling.
- Category
- Writing & Preference
- Introduced
- 2024
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- livebench.ai
Top 15 models
- GPT-5.1 80.2%
- Gemini 2.5 Pro 67.8%
- o1 65.4%
- GPT-4.5 61.5%
- Claude 3.7 Sonnet 59.9%
- Qwen2.5-Max 56.3%
- Claude 3.5 Sonnet 53.8%
- QwQ-32B 51.4%
- Gemini 2.0 Flash (Feb 2025) 51.3%
- o3-mini 50.7%
- Claude 3 Opus 50.4%
- DeepSeek-V3 49.1%
- DeepSeek-R1 48.5%
- GPT-4o 47.6%
- Grok-2 (Dec 2024) 45.6%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Other writing & preference benchmarks
Frequently asked questions
What does LiveBench Language measure?
LiveBench language tasks such as typo fixing and plot unscrambling.
Which model has the highest LiveBench Language score?
As of October 2026, GPT-5.1 has the highest published LiveBench Language score on Noometry at 80.2%, out of 39 models with results.
What is the best open-weight model on LiveBench Language?
QwQ-32B has the highest LiveBench Language accuracy among open-weight models at 51.4%, ranking 8 of 39 overall.