xAI, proprietary

Grok 4.1

Grok 4.1 by xAI ranks 134th of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.5. Its strongest category is agentic & tool use, where it ranks 49th.

Last verified

Specifications

Noometry rank
#134 of 354
Index score
41.5
Evidence
Confirmed 19 results
Provider
xAI
Released
November 17, 2025
Weights
Proprietary
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Grok 4.1 category scores
  1. Coding 33.7
  2. Agentic & Tool Use 34.1
  3. Reasoning 29.5
  4. Math 38.9
  5. Knowledge 39.5
  6. Multilingual 53.4
  7. Instruction Following 73.8
  8. Long Context 43.2
  9. Writing & Preference 62.4
Grok 4.1 category ranks
CategoryScoreRankResults
Coding33.7#2532
Agentic & Tool Use34.1#491
Reasoning29.5#911
Math38.9#1201
Knowledge39.5#1331
Multilingual53.4#681
Instruction Following73.8#1111
Long Context43.2#1001
Writing & Preference62.4#753

Strengths and weaknesses

Categories where Grok 4.1 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Grok 4.1: strongest categories
CategoryScorevs medianRank
Multilingual53.4+6.0#68 of 297, top 23%
Writing & Preference62.4+8.6#75 of 312, top 25%
Reasoning29.5+5.9#91 of 350, top 26%

Weakest categories

Grok 4.1: weakest categories
CategoryScorevs medianRank
Coding33.7−5.0#253 of 340, top 75%
Knowledge39.5+2.2#133 of 314, top 43%
Math38.9+2.4#120 of 327, top 37%

Closest competitors

The models ranked just above and below Grok 4.1. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Grok 4.1
ModelRankScoreBlended $/MSpeed
Granite 4.2 30b#13041.8——Compare
Muse Glimmer#13141.7——Compare
o4-mini#13241.6$1.936Compare
Gemini 3.5 Flash Lite#13341.5$0.85—Compare
GLM-4.6#13541.4$112Compare
Grok 4.1 Fast#13641.4$0.28—Compare
GLM-4.6V#13741.3$0.45—Compare
MiMo-V2-Flash#13841.3$0.18—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Grok 4.1 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena WebDev1214#108 of 113, top 96%thinkingLMArena2026-10-08
LMArena Coding1445#93 of 294, top 32%thinkingLMArena2026-10-08

Agentic & Tool Use

Grok 4.1 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Cybench39%#6 of 21, top 29%Epoch AI

Reasoning

Grok 4.1 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Hard Prompts1435#85 of 297, top 29%LMArena2026-10-08

Math

Grok 4.1 Math benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Math1422#99 of 285, top 35%thinkingLMArena2026-10-08

Knowledge

Grok 4.1 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Expert1417#109 of 273, top 40%thinkingLMArena2026-10-08

Multilingual

Grok 4.1 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1425#68 of 297, top 23%thinkingLMArena2026-10-08
LMArena Chinese1465#81 of 285, top 29%LMArena2026-10-08
LMArena French1448#74 of 223, top 34%thinkingLMArena2026-10-08
LMArena German1446#51 of 231, top 23%LMArena2026-10-08
LMArena Japanese1397#62 of 211, top 30%LMArena2026-10-08
LMArena Korean1407#46 of 213, top 22%LMArena2026-10-08
LMArena Russian1434#61 of 283, top 22%thinkingLMArena2026-10-08
LMArena Spanish1438#69 of 226, top 31%thinkingLMArena2026-10-08

Instruction Following

Grok 4.1 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1400#106 of 298, top 36%LMArena2026-10-08

Long Context

Grok 4.1 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1416#95 of 291, top 33%LMArena2026-10-08

Writing & Preference

Grok 4.1 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1437#70 of 297, top 24%thinkingLMArena2026-10-08
LMArena Creative Writing1411#63 of 295, top 22%LMArena2026-10-08
LMArena Multi-Turn1437#75 of 295, top 26%LMArena2026-10-08

Compare Grok 4.1

Other xAI models

Frequently asked questions

How good is Grok 4.1?

Grok 4.1 by xAI ranks 134th of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.5. Its strongest category is agentic & tool use, where it ranks 49th.

Is Grok 4.1 open source?

No. Grok 4.1 is proprietary and available only through xAI's API and partner platforms.

What are Grok 4.1's strengths and weaknesses?

Relative to other ranked models, Grok 4.1 places best in multilingual, writing & preference, reasoning and lowest in coding, knowledge, math.

What is Grok 4.1 best at?

Its best category is agentic & tool use, where it ranks 49th on Noometry.