Knowledge benchmark
GPQA (HELM) leaderboard
As of October 2026, Gemini 3 Pro has the highest published GPQA (HELM) score on Noometry at 80.3%, out of 57 models with results.
Last verified
About GPQA (HELM)
Graduate-level biology, physics and chemistry questions written to be hard to look up, run by HELM Capabilities with chain-of-thought.
- Category
- Knowledge
- Introduced
- 2023
- Format
- 4-option multiple choice
- Unit
- Percent (random guessing ≈ 25%)
- Official site
- crfm.stanford.edu
Top 15 models
- Gemini 3 Pro 80.3%
- GPT-5 79.2%
- GPT-5 Mini 75.6%
- o3 75.3%
- Gemini 2.5 Pro 74.9%
- o4-mini 73.5%
- Grok 4 72.7%
- Qwen3 235B-A22B 72.7%
- Claude Opus 4 70.8%
- Claude Sonnet 4 70.6%
- Claude Sonnet 4.5 68.6%
- gpt-oss-120b 68.4%
- GPT-5 Nano 67.9%
- Grok-3 mini 67.5%
- DeepSeek-R1 66.6%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Other knowledge benchmarks
- GPQA Diamond
- Humanity's Last Exam
- SimpleQA Verified
- MMLU-Pro
- Confabulations
- Vectara Hallucination Rate
- LMArena Expert
- ARC (AI2) Challenge (reference)
- BoolQ (reference)
- MMLU (reference)
- OpenBookQA (reference)
- TriviaQA (reference)
Frequently asked questions
What does GPQA (HELM) measure?
Graduate-level biology, physics and chemistry questions written to be hard to look up, run by HELM Capabilities with chain-of-thought.
Which model has the highest GPQA (HELM) score?
As of October 2026, Gemini 3 Pro has the highest published GPQA (HELM) score on Noometry at 80.3%, out of 57 models with results.
What is the best open-weight model on GPQA (HELM)?
Qwen3 235B-A22B has the highest GPQA (HELM) accuracy among open-weight models at 72.7%, ranking 8 of 57 overall.