Reasoning benchmark

# Thematic Generalization leaderboard

> Thematic Generalization results for 23 AI models, led by Claude Opus 4.6 at 80.6%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/thematic-generalization
- Last updated: 2026-10-10
- Title: Thematic Generalization Leaderboard (October 2026): Scores by Model

As of October 2026, Claude Opus 4.6 has the highest published Thematic Generalization score on Noometry at 80.6%, out of 23 models with results.

Last verified October 10, 2026

## About Thematic Generalization

Infer a narrow hidden theme from a few examples and anti-examples, then pick the one candidate that fits it among close distractors.

- **Category:** [Reasoning](https://noometry.com/best/reasoning)
- **Introduced:** 2025
- **Format:** Pick the example
- **Unit:** Percent (random guessing ≈ 0%)
- **Official site:** [github.com](https://github.com/lechmazur/generalization)

## Top 15 models

Top models on Thematic Generalization

1.  Claude Opus 4.6 80.6%
2.  GPT-5.4 80%
3.  Gemini 3.1 Pro Preview 79.4%
4.  Claude Sonnet 4.6 76.3%
5.  Claude Opus 4.7 72.8%
6.  GLM-5.1 69.8%
7.  Kimi K2.5 69.4%
8.  Qwen3.5 397B-A17B 65.1%
9.  DeepSeek-V3.2-Exp 65%
10.  Grok 4.20 (Non-Reasoning) 63.8%
11.  Gemini 3.1 Flash Lite 63.3%
12.  GPT-5.4 mini 61.7%
13.  Qwen3.6 Plus 59.5%
14.  Seed 2.0 Pro 57.1%
15.  Gemma 4 31B IT 53%
16.  405060708090

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

Thematic Generalization results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [Claude Opus 4.6](https://noometry.com/models/claude-opus-4-6) | [Anthropic](https://noometry.com/providers/anthropic) | 80.6% | high reasoning | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 2 | [GPT-5.4](https://noometry.com/models/gpt-5-4) | [OpenAI](https://noometry.com/providers/openai) | 80% | xhigh reasoning | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 3 | [Gemini 3.1 Pro Preview](https://noometry.com/models/gemini-3-1-pro-preview) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 79.4% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 4 | [Claude Sonnet 4.6](https://noometry.com/models/claude-sonnet-4-6) | [Anthropic](https://noometry.com/providers/anthropic) | 76.3% | high reasoning | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 5 | [Claude Opus 4.7](https://noometry.com/models/claude-opus-4-7) | [Anthropic](https://noometry.com/providers/anthropic) | 72.8% | high reasoning | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 6 | [GLM-5.1](https://noometry.com/models/glm-5-1) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 69.8% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 7 | [Kimi K2.5](https://noometry.com/models/kimi-k2-5) | [Moonshot AI](https://noometry.com/providers/moonshot) | 69.4% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 8 | [Qwen3.5 397B-A17B](https://noometry.com/models/qwen3-5-397b-a17b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 65.1% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 9 | [DeepSeek-V3.2-Exp](https://noometry.com/models/deepseek-v3-2-exp) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 65% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 10 | [Grok 4.20 (Non-Reasoning)](https://noometry.com/models/grok-4-20) | [xAI](https://noometry.com/providers/xai) | 63.8% | reasoning | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 11 | [Gemini 3.1 Flash Lite](https://noometry.com/models/gemini-3-1-flash-lite) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 63.3% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 12 | [GPT-5.4 mini](https://noometry.com/models/gpt-5-4-mini) | [OpenAI](https://noometry.com/providers/openai) | 61.7% | xhigh reasoning | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 13 | [Qwen3.6 Plus](https://noometry.com/models/qwen3-6-plus) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 59.5% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 14 | [Seed 2.0 Pro](https://noometry.com/models/dola-seed-2-0-pro) |  [![](/logos/bytedance.svg) ByteDance Seed](https://noometry.com/providers/bytedance) | 57.1% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 15 | [Gemma 4 31B IT](https://noometry.com/models/gemma-4-31b-it) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 53% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 16 | [Qwen3.5 122B-A10B](https://noometry.com/models/qwen3-5-122b-a10b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 51.2% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 17 | [MiMo-V2-Pro](https://noometry.com/models/mimo-v2-pro) | [Xiaomi](https://noometry.com/providers/xiaomi) | 45.9% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 18 | [Qwen3.5 27B](https://noometry.com/models/qwen3-5-27b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 45.5% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 19 | [ERNIE 5.0 0110](https://noometry.com/models/ernie-5-0) |  [![](/logos/baidu.svg) Baidu](https://noometry.com/providers/baidu) | 41.7% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 20 | [Trinity Large Thinking](https://noometry.com/models/trinity-large-thinking) |  [![](/logos/arcee.svg) Arcee AI](https://noometry.com/providers/arcee) | 41.6% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 21 | [MiniMax-M2.7](https://noometry.com/models/minimax-m2-7) |  [![](/logos/minimax.svg) MiniMax](https://noometry.com/providers/minimax) | 39.3% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 22 | [Mistral Large 3](https://noometry.com/models/mistral-large-3) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 23% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| 23 | [Mistral Medium 3.1](https://noometry.com/models/mistral-medium-3-1) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 20.3% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |

## Compare the leaders

-   [Claude Opus 4.6 vs GPT-5.4](https://noometry.com/compare/claude-opus-4-6-vs-gpt-5-4)
-   [Claude Opus 4.6 vs Gemini 3.1 Pro Preview](https://noometry.com/compare/claude-opus-4-6-vs-gemini-3-1-pro-preview)
-   [Claude Opus 4.6 vs Claude Sonnet 4.6](https://noometry.com/compare/claude-opus-4-6-vs-claude-sonnet-4-6)
-   [Claude Opus 4.6 vs Claude Opus 4.7](https://noometry.com/compare/claude-opus-4-6-vs-claude-opus-4-7)
-   [GPT-5.4 vs Gemini 3.1 Pro Preview](https://noometry.com/compare/gemini-3-1-pro-preview-vs-gpt-5-4)
-   [GPT-5.4 vs Claude Sonnet 4.6](https://noometry.com/compare/claude-sonnet-4-6-vs-gpt-5-4)

## Other reasoning benchmarks

-   [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2)
-   [SimpleBench](https://noometry.com/benchmarks/simplebench)
-   [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning)
-   [NYT Connections (extended)](https://noometry.com/benchmarks/nyt-connections)
-   [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1)
-   [CritPt](https://noometry.com/benchmarks/critpt)
-   [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles)
-   [EnigmaEval](https://noometry.com/benchmarks/enigmaeval)
-   [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts)
-   [EBR-Bench](https://noometry.com/benchmarks/ebr-bench)
-   [LiveBench Reasoning](https://noometry.com/benchmarks/livebench-reasoning)
-   [Mystery Game Puzzles](https://noometry.com/benchmarks/mystery-game-puzzles)

## Frequently asked questions

### What does Thematic Generalization measure?

Infer a narrow hidden theme from a few examples and anti-examples, then pick the one candidate that fits it among close distractors.

### Which model has the highest Thematic Generalization score?

As of October 2026, Claude Opus 4.6 has the highest published Thematic Generalization score on Noometry at 80.6%, out of 23 models with results.

### What is the best open-weight model on Thematic Generalization?

GLM-5.1 has the highest Thematic Generalization accuracy among open-weight models at 69.8%, ranking 6 of 23 overall.

### Cite this page

Noometry. (2026). Thematic Generalization leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/thematic-generalization

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/thematic-generalization.md).
