Reasoning benchmark
Thematic Generalization leaderboard
As of October 2026, Claude Opus 4.6 has the highest published Thematic Generalization score on Noometry at 80.6%, out of 23 models with results.
Last verified
About Thematic Generalization
Infer a narrow hidden theme from a few examples and anti-examples, then pick the one candidate that fits it among close distractors.
- Category
- Reasoning
- Introduced
- 2025
- Format
- Pick the example
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- github.com
Top 15 models
- Claude Opus 4.6 80.6%
- GPT-5.4 80%
- Gemini 3.1 Pro Preview 79.4%
- Claude Sonnet 4.6 76.3%
- Claude Opus 4.7 72.8%
- GLM-5.1 69.8%
- Kimi K2.5 69.4%
- Qwen3.5 397B-A17B 65.1%
- DeepSeek-V3.2-Exp 65%
- Grok 4.20 (Non-Reasoning) 63.8%
- Gemini 3.1 Flash Lite 63.3%
- GPT-5.4 mini 61.7%
- Qwen3.6 Plus 59.5%
- Seed 2.0 Pro 57.1%
- Gemma 4 31B IT 53%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does Thematic Generalization measure?
Infer a narrow hidden theme from a few examples and anti-examples, then pick the one candidate that fits it among close distractors.
Which model has the highest Thematic Generalization score?
As of October 2026, Claude Opus 4.6 has the highest published Thematic Generalization score on Noometry at 80.6%, out of 23 models with results.
What is the best open-weight model on Thematic Generalization?
GLM-5.1 has the highest Thematic Generalization accuracy among open-weight models at 69.8%, ranking 6 of 23 overall.