Reasoning benchmark

ForecastBench leaderboard

As of October 2026, o3 has the highest published ForecastBench score on Noometry at 62.5, out of 72 models with results.

Last verified

About ForecastBench

Forecasting real-world events that resolve after the model's training cutoff.

Category
Reasoning
Introduced
2024
Format
Probabilistic forecasts
Unit
Raw score
Official site
www.forecastbench.org

Top 15 models

Top models on ForecastBench
  1. o3 62.5
  2. Claude Opus 4.1 62
  3. Claude Sonnet 4.6 62
  4. Claude Sonnet 4.5 61.9
  5. Claude 3.7 Sonnet 61.8
  6. o4-mini 61.8
  7. GPT-4.5 61.7
  8. GPT-4.1 61.5
  9. Claude Haiku 4.5 61.4
  10. GPT-5 61.4
  11. Grok 4.20 (Non-Reasoning) 61.4
  12. MiniMax-M3 61.4
  13. Gemini 2.5 Pro 61.3
  14. Gemini 3 Pro 61.2
  15. Claude Opus 4 61.1

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

ForecastBench results by model
#ModelProviderRatingSettingSourceDate
1o3 OpenAI62.5Epoch AI
2Claude Opus 4.1 Anthropic62Epoch AI
3Claude Sonnet 4.6 Anthropic6216KEpoch AI
4Claude Sonnet 4.5 Anthropic61.9Epoch AI
5Claude 3.7 Sonnet Anthropic61.8Epoch AI
6o4-mini OpenAI61.8Epoch AI
7GPT-4.5 OpenAI61.7Epoch AI
8GPT-4.1 OpenAI61.5Epoch AI
9Claude Haiku 4.5 Anthropic61.4Epoch AI
10GPT-5 OpenAI61.4Epoch AI
11Grok 4.20 (Non-Reasoning) xAI61.4Epoch AI
12MiniMax-M3 MiniMax61.4Epoch AI
13Gemini 2.5 Pro Google61.3Epoch AI
14Gemini 3 Pro Google61.2Epoch AI
15Claude Opus 4 Anthropic61.1Epoch AI
16Claude Sonnet 5 Anthropic61.116KEpoch AI
17Kimi K3 Moonshot AI61.1maxEpoch AI
18GLM-5 Z.ai (Zhipu)61Epoch AI
19GPT-5 Mini OpenAI61Epoch AI
20Grok 4.1 Fast xAI61Epoch AI
21Grok 4 xAI60.9Epoch AI
22Claude 3.5 Sonnet Anthropic60.7Epoch AI
23Claude Opus 4.5 Anthropic60.7Epoch AI
24Gemini 2.5 Flash Google60.6Epoch AI
25GPT-5.5 OpenAI60.6Epoch AI
26Grok 4 Fast xAI60.5Epoch AI
27Claude Opus 4.7 Anthropic60.3Epoch AI
28Grok 4.3 xAI60.3Epoch AI
29Claude Sonnet 4 Anthropic60.2Epoch AI
30Kimi K2 (Jul 2025) Moonshot AI60.2Epoch AI
31GPT-5.2 OpenAI60.1Epoch AI
32Claude Opus 4.6 Anthropic60Epoch AI
33DeepSeek-R1 DeepSeek60Epoch AI
34Claude Opus 4.8 Anthropic59.924KEpoch AI
35Llama 3.1-405B Meta59.9Epoch AI
36Qwen3 235B-A22B Alibaba (Qwen)59.7Epoch AI
37o3-mini OpenAI59.6Epoch AI
38GPT-5.4 OpenAI59.5Epoch AI
39GPT-4 Turbo OpenAI59.4Epoch AI
40GLM-4.5-Air Z.ai (Zhipu)59.2Epoch AI
41DeepSeek-V3 DeepSeek59.1Epoch AI
42GPT-5 Nano OpenAI59.1Epoch AI
43Gemini 3.1 Pro Preview Google59Epoch AI
44Gemini 3.5 Flash Google59Epoch AI
45Llama-3.3-70B-Instruct Meta58.6Epoch AI
46Llama 3-8B Meta58.6Epoch AI
47Gemini 3 Flash Preview Google58.5Epoch AI
48Claude 3 Opus Anthropic58.4Epoch AI
49Gemini 1.5 Pro (May 2024) Google58.4Epoch AI
50QwQ-32B Alibaba (Qwen)58.3Epoch AI
51GPT-5.1 OpenAI58.1Epoch AI
52DeepSeek-V3.1 DeepSeek58Epoch AI
53GPT-4 OpenAI57.8Epoch AI
54GPT-4o OpenAI57.7Epoch AI
55Qwen1.5-110B Alibaba (Qwen)57.7Epoch AI
56Llama 4 Maverick Meta57.5Epoch AI
57Llama 4 Scout Meta57.5Epoch AI
58Qwen2.5 72B Instruct Alibaba (Qwen)57.5Epoch AI
59GPT-5.4 nano OpenAI57.3Epoch AI
60Gemini 2.0 Flash-Lite Google57.1Epoch AI
61Llama 3-70B Meta57.1Epoch AI
62Mistral Large Mistral AI57.1Epoch AI
63GPT-5.4 mini OpenAI57Epoch AI
64Mixtral 8x22B Mistral AI56.3Epoch AI
65Mixtral 8x7B Mistral AI56.3Epoch AI
66DeepSeek V4 Pro DeepSeek56.1Epoch AI
67Gemini 3.1 Flash Lite Google54.4Epoch AI
68Claude 2.1 Anthropic54.2Epoch AI
69Gemini 1.5 Flash (May 2024) Google53.9Epoch AI
70Claude 3 Haiku Anthropic53.2Epoch AI
71Llama 2-70B Meta51.4Epoch AI
72GPT-3.5-turbo OpenAI50.4Epoch AI

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does ForecastBench measure?

Forecasting real-world events that resolve after the model's training cutoff.

Which model has the highest ForecastBench score?

As of October 2026, o3 has the highest published ForecastBench score on Noometry at 62.5, out of 72 models with results.

What is the best open-weight model on ForecastBench?

MiniMax-M3 has the highest ForecastBench score among open-weight models at 61.4, ranking 12 of 72 overall.