Reasoning benchmark
PIQA leaderboard
As of October 2026, GPT-4o mini has the highest published PIQA score on Noometry at 88.7%, out of 27 models with results.
Last verified
About PIQA
Physical commonsense: choose the better way to accomplish a goal.
- Category
- Reasoning
- Introduced
- 2019
- Format
- Binary choice
- Unit
- Percent (random guessing ≈ 50%)
- Official site
- yonatanbisk.com
Top 15 models
- GPT-4o mini 88.7%
- Phi-3.5-MoE 88.6%
- Gemini 1.5 Flash (May 2024) 87.5%
- Llama 3.1-405B 85.9%
- Falcon-180B 84.9%
- DeepSeek-V3 84.7%
- DeepSeek-V2 (MoE-236B, May 2024) 83.9%
- Gemma 2 9B 83.7%
- Mixtral 8x7B 83.6%
- Mistral Nemo 83.5%
- Falcon-40B 83%
- Mistral 7B 83%
- Llama 2-70B 82.8%
- Qwen2.5 72B Instruct 82.6%
- Nemotron-4 15B 82.4%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does PIQA measure?
Physical commonsense: choose the better way to accomplish a goal.
Which model has the highest PIQA score?
As of October 2026, GPT-4o mini has the highest published PIQA score on Noometry at 88.7%, out of 27 models with results.
What is the best open-weight model on PIQA?
Phi-3.5-MoE has the highest PIQA accuracy among open-weight models at 88.6%, ranking 2 of 27 overall.