Reasoning benchmark

PIQA leaderboard

As of October 2026, GPT-4o mini has the highest published PIQA score on Noometry at 88.7%, out of 27 models with results.

Last verified

About PIQA

Physical commonsense: choose the better way to accomplish a goal.

Category
Reasoning
Introduced
2019
Format
Binary choice
Unit
Percent (random guessing ≈ 50%)
Official site
yonatanbisk.com

Top 15 models

Top models on PIQA
  1. GPT-4o mini 88.7%
  2. Phi-3.5-MoE 88.6%
  3. Gemini 1.5 Flash (May 2024) 87.5%
  4. Llama 3.1-405B 85.9%
  5. Falcon-180B 84.9%
  6. DeepSeek-V3 84.7%
  7. DeepSeek-V2 (MoE-236B, May 2024) 83.9%
  8. Gemma 2 9B 83.7%
  9. Mixtral 8x7B 83.6%
  10. Mistral Nemo 83.5%
  11. Falcon-40B 83%
  12. Mistral 7B 83%
  13. Llama 2-70B 82.8%
  14. Qwen2.5 72B Instruct 82.6%
  15. Nemotron-4 15B 82.4%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does PIQA measure?

Physical commonsense: choose the better way to accomplish a goal.

Which model has the highest PIQA score?

As of October 2026, GPT-4o mini has the highest published PIQA score on Noometry at 88.7%, out of 27 models with results.

What is the best open-weight model on PIQA?

Phi-3.5-MoE has the highest PIQA accuracy among open-weight models at 88.6%, ranking 2 of 27 overall.