xAI, proprietary
Grok 4.1
Grok 4.1 by xAI ranks 134th of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.5. Its strongest category is agentic & tool use, where it ranks 49th.
Last verified
Specifications
- Noometry rank
- #134 of 354
- Index score
- 41.5
- Evidence
- Confirmed 19 results
- Provider
- xAI
- Released
- November 17, 2025
- Weights
- Proprietary
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 33.7
- Agentic & Tool Use 34.1
- Reasoning 29.5
- Math 38.9
- Knowledge 39.5
- Multilingual 53.4
- Instruction Following 73.8
- Long Context 43.2
- Writing & Preference 62.4
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 33.7 | #253 | 2 |
| Agentic & Tool Use | 34.1 | #49 | 1 |
| Reasoning | 29.5 | #91 | 1 |
| Math | 38.9 | #120 | 1 |
| Knowledge | 39.5 | #133 | 1 |
| Multilingual | 53.4 | #68 | 1 |
| Instruction Following | 73.8 | #111 | 1 |
| Long Context | 43.2 | #100 | 1 |
| Writing & Preference | 62.4 | #75 | 3 |
Strengths and weaknesses
Categories where Grok 4.1 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multilingual | 53.4 | +6.0 | #68 of 297, top 23% |
| Writing & Preference | 62.4 | +8.6 | #75 of 312, top 25% |
| Reasoning | 29.5 | +5.9 | #91 of 350, top 26% |
Closest competitors
The models ranked just above and below Grok 4.1. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Granite 4.2 30b | #130 | 41.8 | — | — | Compare |
| Muse Glimmer | #131 | 41.7 | — | — | Compare |
| o4-mini | #132 | 41.6 | $1.93 | 6 | Compare |
| Gemini 3.5 Flash Lite | #133 | 41.5 | $0.85 | — | Compare |
| GLM-4.6 | #135 | 41.4 | $1 | 12 | Compare |
| Grok 4.1 Fast | #136 | 41.4 | $0.28 | — | Compare |
| GLM-4.6V | #137 | 41.3 | $0.45 | — | Compare |
| MiMo-V2-Flash | #138 | 41.3 | $0.18 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena WebDev | 1214 | #108 of 113, top 96% | thinking | LMArena | 2026-10-08 |
| LMArena Coding | 1445 | #93 of 294, top 32% | thinking | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Cybench | 39% | #6 of 21, top 29% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Hard Prompts | 1435 | #85 of 297, top 29% | LMArena | 2026-10-08 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Math | 1422 | #99 of 285, top 35% | thinking | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Expert | 1417 | #109 of 273, top 40% | thinking | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1425 | #68 of 297, top 23% | thinking | LMArena | 2026-10-08 |
| LMArena Chinese | 1465 | #81 of 285, top 29% | LMArena | 2026-10-08 | |
| LMArena French | 1448 | #74 of 223, top 34% | thinking | LMArena | 2026-10-08 |
| LMArena German | 1446 | #51 of 231, top 23% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1397 | #62 of 211, top 30% | LMArena | 2026-10-08 | |
| LMArena Korean | 1407 | #46 of 213, top 22% | LMArena | 2026-10-08 | |
| LMArena Russian | 1434 | #61 of 283, top 22% | thinking | LMArena | 2026-10-08 |
| LMArena Spanish | 1438 | #69 of 226, top 31% | thinking | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1400 | #106 of 298, top 36% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1416 | #95 of 291, top 33% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1437 | #70 of 297, top 24% | thinking | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1411 | #63 of 295, top 22% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1437 | #75 of 295, top 26% | LMArena | 2026-10-08 |
Compare Grok 4.1
- Grok 4.1 vs Grok 4
- Grok 4.1 vs Gemini 3.5 Flash Lite
- Grok 4.1 vs GLM-4.6
- Grok 4.1 vs o4-mini
- Grok 4.1 vs Grok 4.1 Fast
- Grok 4.1 vs Muse Glimmer
- Grok 4.1 vs GLM-4.6V
- Grok 4.1 vs GPT-6 Astra
- Grok 4.1 vs Claude Fable 5.1
- Grok 4.1 vs Gemini 3.8 Flash
- Grok 4.1 vs Kimi K3
- Grok 4.1 vs Qwen3.8 Max
- Grok 4.1 vs GLM-5.3
- Grok 4.1 vs Muse Spark 1.3
Other xAI models
- Grok 4.656.9
- Grok 4.555.0
- Grok 4.753.1
- Grok 4.20 (Non-Reasoning)48.6
- Grok 448.1
- Grok 4.20 Multi-Agent46.2
- Grok 4.343.8
- Grok 4.1 Fast41.4
Frequently asked questions
How good is Grok 4.1?
Grok 4.1 by xAI ranks 134th of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.5. Its strongest category is agentic & tool use, where it ranks 49th.
Is Grok 4.1 open source?
No. Grok 4.1 is proprietary and available only through xAI's API and partner platforms.
What are Grok 4.1's strengths and weaknesses?
Relative to other ranked models, Grok 4.1 places best in multilingual, writing & preference, reasoning and lowest in coding, knowledge, math.
What is Grok 4.1 best at?
Its best category is agentic & tool use, where it ranks 49th on Noometry.