Knowledge benchmark

Confabulations leaderboard

As of October 2026, GPT-5 has the lowest Confabulations on Noometry at 10.3%, out of 51 models with results.

Last verified

About Confabulations

Questions about provided documents where the answer is deliberately missing. The score averages the rate of made-up answers and the rate of refusing questions that do have answers. Lower is better.

Category
Knowledge
Introduced
2024
Format
Grounded questions
Unit
Percent, lower is better
Official site
github.com

Top 15 models

Top models on Confabulations
  1. GPT-5 10.3%
  2. Gemini 2.5 Pro 10.6%
  3. Grok-3 mini 10.8%
  4. GLM-4.5 11.3%
  5. o1 11.7%
  6. Qwen3-30B-A3B 12.3%
  7. Grok 4 12.4%
  8. Gemini 2.0 Flash (Feb 2025) 12.4%
  9. DeepSeek-R1 12.7%
  10. Claude Sonnet 4 13.2%
  11. GPT-5 Mini 13.3%
  12. Gemini 1.5 Pro (May 2024) 13.5%
  13. GPT-4.5 13.6%
  14. Grok 3 14.2%
  15. o3-pro 14.2%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Confabulations results by model
#ModelProviderScoreSettingSourceDate
1GPT-5 OpenAI10.3%medium reasoningLech Mazur benchmarks
2Gemini 2.5 Pro Google10.6%Lech Mazur benchmarks
3Grok-3 mini xAI10.8%highLech Mazur benchmarks
4GLM-4.5 Z.ai (Zhipu)11.3%Lech Mazur benchmarks
5o1 OpenAI11.7%medium reasoningLech Mazur benchmarks
6Qwen3-30B-A3B Alibaba (Qwen)12.3%Lech Mazur benchmarks
7Grok 4 xAI12.4%Lech Mazur benchmarks
8Gemini 2.0 Flash (Feb 2025) Google12.4%Lech Mazur benchmarks
9DeepSeek-R1 DeepSeek12.7%Lech Mazur benchmarks
10Claude Sonnet 4 Anthropic13.2%Lech Mazur benchmarks
11GPT-5 Mini OpenAI13.3%medium reasoningLech Mazur benchmarks
12Gemini 1.5 Pro (May 2024) Google13.5%septLech Mazur benchmarks
13GPT-4.5 OpenAI13.6%Lech Mazur benchmarks
14Grok 3 xAI14.2%no reasoningLech Mazur benchmarks
15o3-pro OpenAI14.2%medium reasoningLech Mazur benchmarks
16o3 OpenAI14.4%high reasoningLech Mazur benchmarks
17Claude 3.7 Sonnet Anthropic14.7%Lech Mazur benchmarks
18GPT-4o OpenAI15.3%Lech Mazur benchmarks
19QwQ-32B Alibaba (Qwen)15.6%Lech Mazur benchmarks
20Qwen3 235B-A22B Alibaba (Qwen)15.6%Lech Mazur benchmarks
21gpt-oss-120b OpenAI15.7%medium reasoningLech Mazur benchmarks
22o4-mini OpenAI15.8%high reasoningLech Mazur benchmarks
23Claude Opus 4 Anthropic15.9%Lech Mazur benchmarks
24Chatgpt 4o Latest 20250326 OpenAI16.6%Lech Mazur benchmarks
25Gemini 2.5 Flash Google16.8%Lech Mazur benchmarks
26Claude Opus 4.1 Anthropic17.1%Lech Mazur benchmarks
27Llama 3.1-405B Meta17.6%Lech Mazur benchmarks
28o3-mini OpenAI17.9%medium reasoningLech Mazur benchmarks
29Gemini 2.0 Pro Google18.4%Lech Mazur benchmarks
30o1-mini OpenAI18.6%Lech Mazur benchmarks
31Qwen2.5 72B Instruct Alibaba (Qwen)19.1%Lech Mazur benchmarks
32Claude 3.5 Sonnet Anthropic19.9%Lech Mazur benchmarks
33Grok-2 (Dec 2024) xAI20.1%Lech Mazur benchmarks
34Kimi K2 (Jul 2025) Moonshot AI20.4%Lech Mazur benchmarks
35Mistral Large Mistral AI21.4%Lech Mazur benchmarks
36Qwen2.5-Max Alibaba (Qwen)21.8%Lech Mazur benchmarks
37Mistral Medium 3 Mistral AI21.9%Lech Mazur benchmarks
38Llama 4 Maverick Meta22.6%Lech Mazur benchmarks
39Claude 3 Opus Anthropic22.7%Lech Mazur benchmarks
40Llama-3.3-70B-Instruct Meta22.8%Lech Mazur benchmarks
41MiniMax-01 MiniMax23.9%Lech Mazur benchmarks
42Mistral Small 3 Mistral AI25.2%Lech Mazur benchmarks
43DeepSeek-V3 DeepSeek26.1%Lech Mazur benchmarks
44Gemma 2 27B Google27.1%Lech Mazur benchmarks
45GPT-4 Turbo OpenAI28.4%Lech Mazur benchmarks
46Phi-4 Microsoft29.4%Lech Mazur benchmarks
47Amazon Nova Pro Amazon30.1%Lech Mazur benchmarks
48Claude 3 Haiku Anthropic34.2%Lech Mazur benchmarks
49Claude 3.5 Haiku Anthropic36.7%Lech Mazur benchmarks
50GPT-4o mini OpenAI37.2%Lech Mazur benchmarks
51Gemma 3 27B Google40.3%Lech Mazur benchmarks

Compare the leaders

Other knowledge benchmarks

Frequently asked questions

What does Confabulations measure?

Questions about provided documents where the answer is deliberately missing. The score averages the rate of made-up answers and the rate of refusing questions that do have answers. Lower is better.

Which model has the highest Confabulations score?

As of October 2026, GPT-5 has the lowest Confabulations on Noometry at 10.3%, out of 51 models with results.

What is the best open-weight model on Confabulations?

GLM-4.5 has the lowest Confabulations result among open-weight models at 11.3%, ranking 4 of 51 overall.