Knowledge benchmark
ARC (AI2) Challenge leaderboard
As of October 2026, DeepSeek-V3 has the highest published ARC (AI2) Challenge score on Noometry at 95.3%, out of 39 models with results.
Last verified
About ARC (AI2) Challenge
Grade-school science multiple-choice questions that retrieval alone cannot answer.
- Category
- Knowledge
- Introduced
- 2018
- Format
- Multiple choice
- Unit
- Percent (random guessing ≈ 25%)
- Official site
- allenai.org
Top 15 models
- DeepSeek-V3 95.3%
- Llama 3.1-405B 95.3%
- Qwen2.5 72B Instruct 94.5%
- DeepSeek-V2 (MoE-236B, May 2024) 92.2%
- phi-3-medium 14B 91.6%
- Phi 3 Small 8k Instruct 90.7%
- GPT-3.5-turbo 87.4%
- Mixtral 8x7B 87.3%
- Claude Instant 86.3%
- Phi 3 Mini 4k Instruct 84.9%
- Qwen-14B 84.4%
- Llama 3-8B 82.8%
- Mistral 7B 78.6%
- Gemma 7B 78.3%
- Llama 2-70B 78.3%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Other knowledge benchmarks
- GPQA Diamond
- Humanity's Last Exam
- SimpleQA Verified
- MMLU-Pro
- Confabulations
- Vectara Hallucination Rate
- LMArena Expert
- GPQA (HELM)
- BoolQ (reference)
- MMLU (reference)
- OpenBookQA (reference)
- TriviaQA (reference)
Frequently asked questions
What does ARC (AI2) Challenge measure?
Grade-school science multiple-choice questions that retrieval alone cannot answer.
Which model has the highest ARC (AI2) Challenge score?
As of October 2026, DeepSeek-V3 has the highest published ARC (AI2) Challenge score on Noometry at 95.3%, out of 39 models with results.
What is the best open-weight model on ARC (AI2) Challenge?
DeepSeek-V3 has the highest ARC (AI2) Challenge accuracy among open-weight models at 95.3%, ranking 1 of 39 overall.