Knowledge benchmark
Confabulations leaderboard
As of October 2026, GPT-5 has the lowest Confabulations on Noometry at 10.3%, out of 51 models with results.
Last verified
About Confabulations
Questions about provided documents where the answer is deliberately missing. The score averages the rate of made-up answers and the rate of refusing questions that do have answers. Lower is better.
- Category
- Knowledge
- Introduced
- 2024
- Format
- Grounded questions
- Unit
- Percent, lower is better
- Official site
- github.com
Top 15 models
- GPT-5 10.3%
- Gemini 2.5 Pro 10.6%
- Grok-3 mini 10.8%
- GLM-4.5 11.3%
- o1 11.7%
- Qwen3-30B-A3B 12.3%
- Grok 4 12.4%
- Gemini 2.0 Flash (Feb 2025) 12.4%
- DeepSeek-R1 12.7%
- Claude Sonnet 4 13.2%
- GPT-5 Mini 13.3%
- Gemini 1.5 Pro (May 2024) 13.5%
- GPT-4.5 13.6%
- Grok 3 14.2%
- o3-pro 14.2%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Other knowledge benchmarks
- GPQA Diamond
- Humanity's Last Exam
- SimpleQA Verified
- MMLU-Pro
- Vectara Hallucination Rate
- LMArena Expert
- GPQA (HELM)
- ARC (AI2) Challenge (reference)
- BoolQ (reference)
- MMLU (reference)
- OpenBookQA (reference)
- TriviaQA (reference)
Frequently asked questions
What does Confabulations measure?
Questions about provided documents where the answer is deliberately missing. The score averages the rate of made-up answers and the rate of refusing questions that do have answers. Lower is better.
Which model has the highest Confabulations score?
As of October 2026, GPT-5 has the lowest Confabulations on Noometry at 10.3%, out of 51 models with results.
What is the best open-weight model on Confabulations?
GLM-4.5 has the lowest Confabulations result among open-weight models at 11.3%, ranking 4 of 51 overall.