Knowledge benchmark
TriviaQA leaderboard
As of October 2026, Llama 2-70B has the highest published TriviaQA score on Noometry at 87.6%, out of 25 models with results.
Last verified
About TriviaQA
Trivia questions paired with evidence documents.
- Category
- Knowledge
- Introduced
- 2017
- Format
- Short answer
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- nlp.cs.washington.edu
Top 15 models
- Llama 2-70B 87.6%
- Claude 2 87.5%
- Claude 1.3 86.7%
- GPT-3.5-turbo 85.8%
- GPT-4 84.8%
- Llama 2-34B 84.6%
- DeepSeek-V3 82.9%
- Llama 3.1-405B 82.7%
- Mixtral 8x7B 82.2%
- DeepSeek-V2 (MoE-236B, May 2024) 80%
- Falcon-40B 79.9%
- Llama 2-13B 79.6%
- Claude Instant 78.9%
- Llama 13b 77.9%
- Mistral 7B 75.2%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | Llama 2-70B | 87.6% | Epoch AI | |||
| 2 | Claude 2 | Anthropic | 87.5% | Epoch AI | ||
| 3 | Claude 1.3 | Anthropic | 86.7% | Epoch AI | ||
| 4 | GPT-3.5-turbo | OpenAI | 85.8% | Epoch AI | ||
| 5 | GPT-4 | OpenAI | 84.8% | Epoch AI | ||
| 6 | Llama 2-34B | 84.6% | Epoch AI | |||
| 7 | DeepSeek-V3 | 82.9% | Epoch AI | |||
| 8 | Llama 3.1-405B | 82.7% | Epoch AI | |||
| 9 | Mixtral 8x7B | 82.2% | Epoch AI | |||
| 10 | DeepSeek-V2 (MoE-236B, May 2024) | 80% | Epoch AI | |||
| 11 | Falcon-40B | 79.9% | Epoch AI | |||
| 12 | Llama 2-13B | 79.6% | Epoch AI | |||
| 13 | Claude Instant | Anthropic | 78.9% | Epoch AI | ||
| 14 | Llama 13b | 77.9% | Epoch AI | |||
| 15 | Mistral 7B | 75.2% | Epoch AI | |||
| 16 | phi-3-medium 14B | 73.9% | Epoch AI | |||
| 17 | Llama 2-7B | 73.7% | Epoch AI | |||
| 18 | Gemma 7B | 72.3% | Epoch AI | |||
| 19 | Qwen2.5 72B Instruct | 71.9% | Epoch AI | |||
| 20 | Llama 3-8B | 67.7% | Epoch AI | |||
| 21 | Falcon-7B | 64.6% | Epoch AI | |||
| 22 | Phi 3 Mini 4k Instruct | 64% | Epoch AI | |||
| 23 | Phi 3 Small 8k Instruct | 58.1% | Epoch AI | |||
| 24 | Gemma 2B | 53.2% | Epoch AI | |||
| 25 | Phi-2 | 45.2% | Epoch AI |
Compare the leaders
Other knowledge benchmarks
- GPQA Diamond
- Humanity's Last Exam
- SimpleQA Verified
- MMLU-Pro
- Confabulations
- Vectara Hallucination Rate
- LMArena Expert
- GPQA (HELM)
- ARC (AI2) Challenge (reference)
- BoolQ (reference)
- MMLU (reference)
- OpenBookQA (reference)
Frequently asked questions
What does TriviaQA measure?
Trivia questions paired with evidence documents.
Which model has the highest TriviaQA score?
As of October 2026, Llama 2-70B has the highest published TriviaQA score on Noometry at 87.6%, out of 25 models with results.
What is the best open-weight model on TriviaQA?
Llama 2-70B has the highest TriviaQA accuracy among open-weight models at 87.6%, ranking 1 of 25 overall.