Multimodal benchmark
ScienceQA leaderboard
As of October 2026, GPT-4o has the highest published ScienceQA score on Noometry at 88.5%, out of 6 models with results.
Last verified
About ScienceQA
Multimodal grade-school science questions with image and text context.
- Category
- Multimodal
- Introduced
- 2022
- Format
- Multiple choice
- Unit
- Percent (random guessing ≈ 25%)
- Official site
- scienceqa.github.io
Top 6 models
- GPT-4o 88.5%
- Gemini 1.0 Pro Vision 79.7%
- Claude 3 Haiku 72%
- Llama 2-13B 55.8%
- Llama 13b 43.3%
- Llama 2-7B 43.1%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | GPT-4o | OpenAI | 88.5% | Epoch AI | ||
| 2 | Gemini 1.0 Pro Vision | 79.7% | Epoch AI | |||
| 3 | Claude 3 Haiku | Anthropic | 72% | Epoch AI | ||
| 4 | Llama 2-13B | 55.8% | Epoch AI | |||
| 5 | Llama 13b | 43.3% | Epoch AI | |||
| 6 | Llama 2-7B | 43.1% | Epoch AI |
Compare the leaders
Other multimodal benchmarks
- LMArena Vision
- Video-MME
- GeoBench
- VPCT
- Blueprint-Bench 2
- Furniture Assembly
- LMArena Document (reference)
- MindCube (reference)
- SpatialViz-Bench (reference)
Frequently asked questions
What does ScienceQA measure?
Multimodal grade-school science questions with image and text context.
Which model has the highest ScienceQA score?
As of October 2026, GPT-4o has the highest published ScienceQA score on Noometry at 88.5%, out of 6 models with results.
What is the best open-weight model on ScienceQA?
Llama 2-13B has the highest ScienceQA accuracy among open-weight models at 55.8%, ranking 4 of 6 overall.