Reasoning benchmark

BIG-Bench Hard leaderboard

As of October 2026, Gemini 1.5 Pro (May 2024) has the highest published BIG-Bench Hard score on Noometry at 89.2%, out of 27 models with results.

Last verified

About BIG-Bench Hard

23 challenging BIG-Bench tasks where early models fell short of human raters.

Category
Reasoning
Introduced
2022
Format
Mixed
Unit
Percent (random guessing ≈ 25%)
Official site
github.com

Top 15 models

Top models on BIG-Bench Hard
  1. Gemini 1.5 Pro (May 2024) 89.2%
  2. DeepSeek-V3 87.5%
  3. Llama 3.1-405B 82.9%
  4. phi-3-medium 14B 81.4%
  5. Qwen2.5 72B Instruct 79.8%
  6. Phi 3 Small 8k Instruct 79.1%
  7. DeepSeek-V2 (MoE-236B, May 2024) 78.8%
  8. GPT-4 75.1%
  9. Phi 3 Mini 4k Instruct 71.7%
  10. Yi-34B 71.7%
  11. Llama 2-70B 64.9%
  12. GPT-3.5-turbo 61.6%
  13. Phi-2 59.4%
  14. Nemotron-4 15B 58.7%
  15. Llama 2-13B 58.2%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does BIG-Bench Hard measure?

23 challenging BIG-Bench tasks where early models fell short of human raters.

Which model has the highest BIG-Bench Hard score?

As of October 2026, Gemini 1.5 Pro (May 2024) has the highest published BIG-Bench Hard score on Noometry at 89.2%, out of 27 models with results.

What is the best open-weight model on BIG-Bench Hard?

DeepSeek-V3 has the highest BIG-Bench Hard accuracy among open-weight models at 87.5%, ranking 2 of 27 overall.