Best model for your job
Best AI model for hard reasoning
For reasoning, GPT-6 Astra scores highest on the current data (87.6), weighting reasoning 70%, math 30%. The cheapest model in the top 10 is GPT-6.1 Sol at $2 / $10 per million tokens.
Last verified
Weights: Reasoning 70%, Math 30%. Novel problem solving is measured best by abstraction puzzles and math.
| # | Model | Provider | Fit | Reasoning | Math | Input $/M | Output $/M |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | OpenAI | 87.6 | 85.1 | 93.5 | $10 | $50 |
| 2 | GPT-6.1 Sol | OpenAI | 85.4 | 81.9 | 93.7 | $2 | $10 |
| 3 | Claude Opus 5.5 | Anthropic | 83.7 | 80.2 | 91.8 | $4 | $20 |
| 4 | Claude Fable 5.1 | Anthropic | 80.6 | 76.7 | 89.6 | $10 | $50 |
| 5 | Claude Fable 5 | Anthropic | 80.3 | 76.8 | 88.5 | $10 | $50 |
| 6 | Claude Opus 5 | Anthropic | 79.9 | 77.2 | 86.2 | $5 | $25 |
| 7 | GPT-5.6 Sol | OpenAI | 78.0 | 74.8 | 85.6 | $4 | $20 |
| 8 | GPT-6 Sol | OpenAI | 78.0 | 74.0 | 87.2 | $2 | $10 |
| 9 | GPT-5.5 Pro | OpenAI | 76.5 | 73.3 | 84.0 | $30 | $180 |
| 10 | GPT-5.5 | OpenAI | 75.4 | 72.8 | 81.7 | $5 | $30 |
| 11 | Gemini 3.8 Flash | 73.4 | 76.9 | 65.3 | $0.75 | $3.75 | |
| 12 | GPT-5.4 Pro | OpenAI | 71.2 | 70.7 | 72.4 | $30 | $180 |
| 13 | Gemini 3.7 Flash | 69.9 | 70.0 | 69.6 | $0.75 | $3.75 | |
| 14 | Gemini 3.1 Pro Preview | 68.8 | 71.7 | 62.1 | $2 | $12 | |
| 15 | Claude Opus 4.8 | Anthropic | 68.8 | 64.7 | 78.4 | $5 | $25 |
| 16 | GPT-5.6 Terra | OpenAI | 66.9 | 60.7 | 81.6 | $2 | $12 |
| 17 | Kimi K3 (open weights) | Moonshot AI | 66.4 | 63.0 | 74.2 | $3 | $15 |
| 18 | GPT-5.4 | OpenAI | 65.3 | 61.8 | 73.5 | $2.50 | $15 |
| 19 | Claude Sonnet 5.5 | Anthropic | 64.2 | 54.0 | 87.9 | $2 | $10 |
| 20 | Grok 4.6 | xAI | 63.1 | 61.4 | 67.0 | $2 | $6 |
| 21 | Gemini 3.5 Flash | 62.2 | 62.8 | 60.7 | $1.50 | $9 | |
| 22 | Qwen3.8 Max | Alibaba (Qwen) | 60.0 | 54.4 | 73.2 | $2 | $6 |
| 23 | Muse Spark 1.3 | Meta | 59.8 | 54.0 | 73.1 | $1.25 | $4.25 |
| 24 | Claude Opus 4.6 | Anthropic | 59.3 | 57.8 | 63.0 | $5 | $25 |
| 25 | DeepSeek V4 Pro (open weights) | DeepSeek | 59.0 | 56.5 | 64.8 | $0.66 | $1.98 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Top two head to head: GPT-6 Astra vs GPT-6.1 Sol
Frequently asked questions
What is the best ai model for hard reasoning?
For reasoning, GPT-6 Astra scores highest on the current data (87.6), weighting reasoning 70%, math 30%. The cheapest model in the top 10 is GPT-6.1 Sol at $2 / $10 per million tokens.
How is this shortlist built?
Novel problem solving is measured best by abstraction puzzles and math. Each model's category scores are blended with those weights; only ranked models with results in every needed category are listed.
Other jobs
- Best AI model for coding
- Best AI model for debugging
- Best AI model for building websites
- Best local LLM for coding
- Best AI model for agents
- Best AI model for RAG
- Best AI model for research
- Best AI model for studying
- Best AI model for math
- Best AI model for data analysis
- Best AI model for Excel
- Best AI model for data extraction
- Best AI model for writing
- Best AI model for emails
- Best AI model for copywriting
- Best AI model for creative writing
- Best AI model for translation
- Best AI model for language learning
- Best AI model for customer service
- Best AI model for presentations
- Best AI model for PDFs
- Best AI assistant for everyday use