Writing & Preference benchmark
WildBench leaderboard
As of October 2026, Qwen3 235B-A22B has the highest published WildBench score on Noometry at 86.6%, out of 57 models with results.
Last verified
About WildBench
Challenging requests taken from real chatbot conversations, graded by an LLM judge with task-specific checklists. HELM Capabilities run.
- Category
- Writing & Preference
- Introduced
- 2024
- Size
- 1,024 tasks
- Format
- LLM-judged answers
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- crfm.stanford.edu
Top 15 models
- Qwen3 235B-A22B 86.6%
- GPT-5.1 86.3%
- Kimi K2 (Jul 2025) 86.2%
- o3 86.1%
- Gemini 3 Pro 85.9%
- GPT-5 85.7%
- Gemini 2.5 Pro 85.7%
- GPT-5 Mini 85.5%
- Claude Sonnet 4.5 85.4%
- GPT-4.1 85.4%
- o4-mini 85.4%
- Claude Opus 4 85.2%
- Grok 3 84.9%
- gpt-oss-120b 84.5%
- Claude Haiku 4.5 83.9%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Other writing & preference benchmarks
Frequently asked questions
What does WildBench measure?
Challenging requests taken from real chatbot conversations, graded by an LLM judge with task-specific checklists. HELM Capabilities run.
Which model has the highest WildBench score?
As of October 2026, Qwen3 235B-A22B has the highest published WildBench score on Noometry at 86.6%, out of 57 models with results.
What is the best open-weight model on WildBench?
Qwen3 235B-A22B has the highest WildBench accuracy among open-weight models at 86.6%, ranking 1 of 57 overall.