xAI, proprietary
Grok 4.6
Grok 4.6 by xAI ranks 21st of 354 ranked models on the Noometry Index as of October 2026, with a score of 56.9. Its strongest category is coding, where it ranks 16th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 500K-token context window.
Last verified
Specifications
- Noometry rank
- #21 of 354
- Index score
- 56.9
- Evidence
- Confirmed 49 results
- Provider
- xAI
- Released
- August 12, 2026
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 500K
- Max output
- 500K
- Input price
- $2 / M
- Output price
- $6 / M
- Blended price
- $3 / M
- Output speed
- Not measured
- Value
- #153 of 219
- Knowledge cutoff
- February 2026
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 58.5
- Agentic & Tool Use 39.4
- Reasoning 61.4
- Math 67.0
- Knowledge 63.3
- Multimodal 43.6
- Multilingual 53.0
- Instruction Following 75.4
- Long Context 44.5
- Writing & Preference 62.3
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 58.5 | #16 | 8 |
| Agentic & Tool Use | 39.4 | #27 | 2 |
| Reasoning | 61.4 | #20 | 11 |
| Math | 67.0 | #24 | 5 |
| Knowledge | 63.3 | #20 | 3 |
| Multimodal | 43.6 | #23 | 3 |
| Multilingual | 53.0 | #74 | 1 |
| Instruction Following | 75.4 | #63 | 1 |
| Long Context | 44.5 | #66 | 1 |
| Writing & Preference | 62.3 | #80 | 3 |
Strengths and weaknesses
Categories where Grok 4.6 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Writing & Preference | 62.3 | +8.5 | #80 of 312, top 26% |
| Multilingual | 53.0 | +5.6 | #74 of 297, top 25% |
| Long Context | 44.5 | +3.6 | #66 of 296, top 23% |
Closest competitors
The models ranked just above and below Grok 4.6. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| GPT-5.6 Terra | #17 | 59.2 | $4.50 | 11 | Compare |
| GPT-5.4 Pro | #18 | 58.9 | $67.50 | — | Compare |
| Claude Opus 4.7 | #19 | 58.3 | $10 | 33 | Compare |
| Claude Opus 4.6 | #20 | 58.2 | $10 | 19 | Compare |
| Qwen3.8 Max | #22 | 56.8 | $3 | — | Compare |
| Gemini 3.1 Pro Preview | #23 | 56.7 | $4.50 | — | Compare |
| Gemini 4 Argon | #24 | 56.5 | — | — | Compare |
| Grok 4.5 | #25 | 55.0 | $3 | 4 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| DeepSWE | 65.2% | high | Epoch AI | ||
| DeepSWE | 41.6% | low | Epoch AI | ||
| DeepSWE | 67.5% | #11 of 29, top 38% | medium | Epoch AI | |
| DeepSWE | 66.7% | xhigh | Epoch AI | ||
| FrontierCode | 48% | #9 of 37, top 25% | Epoch AI | ||
| CursorBench | 40.4% | high | Epoch AI | ||
| CursorBench | 33.4% | low | Epoch AI | ||
| CursorBench | 36.1% | medium | Epoch AI | ||
| CursorBench | 41.4% | #9 of 14, top 65% | xhigh | Epoch AI | |
| LMArena WebDev | 1617 | #20 of 113, top 18% | high | LMArena | 2026-10-08 |
| FrontierSWE | 25.3% | #12 of 18, top 67% | xhigh | Epoch AI | |
| SciCode | 56.5% | #20 of 121, top 17% | high | Epoch AI | |
| SciCode | 48.4% | low | Epoch AI | ||
| SciCode | 54.6% | medium | Epoch AI | ||
| SciCode | 51.6% | xhigh | Epoch AI | ||
| WeirdML | 67.3% | #21 of 119, top 18% | high | Epoch AI | |
| LMArena Coding | 1465 | #66 of 294, top 23% | high | LMArena | 2026-10-08 |
| ALE-Bench | 1,508 | #17 of 105, top 17% | xhigh | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| APEX-Agents | 65.3% | #7 of 49, top 15% | Epoch AI | ||
| GDP.pdf | 16% | high | Epoch AI | ||
| GDP.pdf | 17.2% | #23 of 36, top 64% | xhigh | Epoch AI | |
| Vending-Bench 2 | 9,047 | #9 of 60, top 15% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 65.1% | high | Epoch AI | ||
| ARC-AGI-2 | 27.6% | low | Epoch AI | ||
| ARC-AGI-2 | 61.3% | medium | Epoch AI | ||
| ARC-AGI-2 | 67.1% | #22 of 83, top 27% | xhigh | Epoch AI | |
| SimpleBench | 75.9% | #7 of 77, top 10% | Epoch AI | ||
| NYT Connections (extended) | 80% | #35 of 91, top 39% | xhigh reasoning | Lech Mazur benchmarks | |
| ARC-AGI-1 | 87% | high | Epoch AI | ||
| ARC-AGI-1 | 74.8% | low | Epoch AI | ||
| ARC-AGI-1 | 87.5% | #30 of 83, top 37% | medium | Epoch AI | |
| ARC-AGI-1 | 87% | xhigh | Epoch AI | ||
| CritPt | 17.1% | high | Epoch AI | ||
| CritPt | 5.7% | low | Epoch AI | ||
| CritPt | 17.7% | medium | Epoch AI | ||
| CritPt | 19.7% | #25 of 134, top 19% | xhigh | Epoch AI | |
| Chess Puzzles | 40% | #21 of 129, top 17% | high | Epoch AI | 2026-08-12 |
| Chess Puzzles | 31% | xhigh | Epoch AI | 2026-08-14 | |
| EBR-Bench | 30.5% | #10 of 24, top 42% | xhigh | Epoch AI | 2026-08-18 |
| LMArena Hard Prompts | 1447 | #68 of 297, top 23% | high | LMArena | 2026-10-08 |
| Mystery Game Puzzles | 34% | #22 of 74, top 30% | xhigh | Epoch AI | 2026-08-14 |
| DTBench | 97.3% | #7 of 151, top 5% | xhigh | Epoch AI | |
| LMCA | 48.5% | #25 of 125, top 20% | xhigh | Epoch AI | |
| Epoch Capabilities Index | 156.44 | #21 of 213, top 10% | Epoch AI | 2026-08-12 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 66% | #30 of 81, top 38% | xhigh | Epoch AI | 2026-08-14 |
| FrontierMath Tier 4 | 31.7% | #27 of 63, top 43% | xhigh | Epoch AI | 2026-08-14 |
| OTIS Mock AIME 2024-2025 | 97.8% | high | Epoch AI | 2026-08-12 | |
| OTIS Mock AIME 2024-2025 | 99.2% | #13 of 173, top 8% | xhigh | Epoch AI | 2026-08-14 |
| ProofBench | 51% | #27 of 77, top 36% | Epoch AI | ||
| LMArena Math | 1423 | #96 of 285, top 34% | high | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 94% | #11 of 186, top 6% | high | Epoch AI | 2026-08-12 |
| GPQA Diamond | 93.2% | xhigh | Epoch AI | 2026-08-14 | |
| SimpleQA Verified | 49.3% | #27 of 77, top 36% | high | Epoch AI | 2026-08-27 |
| SimpleQA Verified | 48.9% | xhigh | Epoch AI | 2026-08-27 | |
| LMArena Expert | 1467 | #52 of 273, top 20% | high | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1263 | #44 of 122, top 37% | high | LMArena | 2026-10-09 |
| Blueprint-Bench 2 | 33.2% | #11 of 31, top 36% | Epoch AI | ||
| Furniture Assembly | 40% | #15 of 31, top 49% | xhigh | Epoch AI | 2026-09-24 |
| LMArena Document | 1452 | #20 of 38, top 53% | high | LMArena | 2026-09-13 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1420 | #74 of 297, top 25% | high | LMArena | 2026-10-08 |
| LMArena Chinese | 1480 | #64 of 285, top 23% | high | LMArena | 2026-10-08 |
| LMArena French | 1461 | #47 of 223, top 22% | high | LMArena | 2026-10-08 |
| LMArena German | 1431 | #64 of 231, top 28% | high | LMArena | 2026-10-08 |
| LMArena Japanese | 1376 | #83 of 211, top 40% | high | LMArena | 2026-10-08 |
| LMArena Korean | 1397 | #57 of 213, top 27% | high | LMArena | 2026-10-08 |
| LMArena Russian | 1422 | #79 of 283, top 28% | high | LMArena | 2026-10-08 |
| LMArena Spanish | 1404 | #106 of 226, top 47% | high | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1431 | #59 of 298, top 20% | high | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1454 | #49 of 291, top 17% | high | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1428 | #80 of 297, top 27% | high | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1428 | #49 of 295, top 17% | high | LMArena | 2026-10-08 |
| LMArena Multi-Turn | 1425 | #90 of 295, top 31% | high | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $2 | $6 | $0.50 | 2026-10-10 |
| bedrock | $2 | $6 | $0.50 | 2026-10-10 |
| openrouter | $2 | $6 | $0.50 | 2026-10-10 |
| vertex | $2 | $6 | $0.50 | 2026-10-10 |
| xai | $2 | $6 | $0.50 | 2026-10-10 |
Compare Grok 4.6
- Grok 4.6 vs Grok 4.5
- Grok 4.6 vs Claude Opus 4.6
- Grok 4.6 vs Qwen3.8 Max
- Grok 4.6 vs Claude Opus 4.7
- Grok 4.6 vs Gemini 3.1 Pro Preview
- Grok 4.6 vs GPT-5.4 Pro
- Grok 4.6 vs Gemini 4 Argon
- Grok 4.6 vs GPT-6 Astra
- Grok 4.6 vs Claude Fable 5.1
- Grok 4.6 vs Gemini 3.8 Flash
- Grok 4.6 vs Kimi K3
- Grok 4.6 vs GLM-5.3
- Grok 4.6 vs Muse Spark 1.3
- Grok 4.6 vs DeepSeek V4 Pro
Other xAI models
- Grok 4.555.0
- Grok 4.753.1
- Grok 4.20 (Non-Reasoning)48.6
- Grok 448.1
- Grok 4.20 Multi-Agent46.2
- Grok 4.343.8
- Grok 4.141.5
- Grok 4.1 Fast41.4
Frequently asked questions
How good is Grok 4.6?
Grok 4.6 by xAI ranks 21st of 354 ranked models on the Noometry Index as of October 2026, with a score of 56.9. Its strongest category is coding, where it ranks 16th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 500K-token context window.
How much does Grok 4.6 cost?
Grok 4.6 costs $2 per million input tokens and $6 per million output tokens on xAI's own API, with cached input at $0.50.
What is Grok 4.6's context window?
Grok 4.6 accepts up to 500K tokens of input and can write up to 500K tokens in one response.
Is Grok 4.6 open source?
No. Grok 4.6 is proprietary and available only through xAI's API and partner platforms.
What are Grok 4.6's strengths and weaknesses?
Relative to other ranked models, Grok 4.6 places best in coding, reasoning, knowledge and lowest in writing & preference, multilingual, long context.
What is Grok 4.6 best at?
Its best category is coding, where it ranks 16th on Noometry.