xAI, proprietary

Grok 4

Grok 4 by xAI ranks 56th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 4th.

Last verified

Specifications

Noometry rank
#56 of 354
Index score
48.1
Evidence
Confirmed 48 results
Provider
xAI
Released
July 9, 2025
Weights
Proprietary
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
1 tokens/s Kagi
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Grok 4 category scores
  1. Coding 50.3
  2. Agentic & Tool Use 32.3
  3. Reasoning 36.7
  4. Math 48.4
  5. Knowledge 53.8
  6. Multimodal 33.7
  7. Multilingual 51.8
  8. Instruction Following 79.2
  9. Long Context 63.1
  10. Writing & Preference 58.5
Grok 4 category ranks
CategoryScoreRankResults
Coding50.3#463
Agentic & Tool Use32.3#686
Reasoning36.7#656
Math48.4#643
Knowledge53.8#555
Multimodal33.7#942
Multilingual51.8#1031
Instruction Following79.2#52
Long Context63.1#42
Writing & Preference58.5#1165

Strengths and weaknesses

Categories where Grok 4 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Grok 4: strongest categories
CategoryScorevs medianRank
Long Context63.1+22.2#4 of 296, top 2%
Instruction Following79.2+8.0#5 of 305, top 2%
Coding50.3+11.6#46 of 340, top 14%

Weakest categories

Grok 4: weakest categories
CategoryScorevs medianRank
Multimodal33.7−4.8#94 of 128, top 74%
Agentic & Tool Use32.3+1.9#68 of 154, top 45%
Writing & Preference58.5+4.7#116 of 312, top 38%

Closest competitors

The models ranked just above and below Grok 4. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Grok 4
ModelRankScoreBlended $/MSpeed
Claude Haiku 5.5#5249.5$0.20—Compare
GPT-5.1#5349.0$3.44—Compare
Grok 4.20 (Non-Reasoning)#5448.6$1.5661Compare
MiMo-V2.6-Flash#5548.5$0.18—Compare
Kimi K2.5#5748.1$0.9066Compare
Step 5 Preview#5847.9$1.43—Compare
GLM-5.1#5947.8$2.15—Compare
Kimi K2.6#6047.7$1.71—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Grok 4 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot79.6%#5 of 44, top 12%Epoch AI
Aider Polyglot79.6%highEpoch AI
WeirdML45.7%#59 of 119, top 50%Epoch AI
LMArena Coding1408#130 of 294, top 45%LMArena2026-10-08

Agentic & Tool Use

Grok 4 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench27.2%#33 of 41, top 81%Epoch AI
Berkeley Function Calling Leaderboard63%#8 of 49, top 17%promptBerkeley Function Calling Leaderboard
GDPval21.1%#10 of 11, top 91%highEpoch AI
Cybench43%#4 of 21, top 20%Epoch AI
DeepResearch Bench47.3%#11 of 24, top 46%Epoch AI
BALROG43.6%#9 of 35, top 26%Epoch AI
LMArena Search1142#27 of 32, top 85%LMArena2026-08-24
METR Time Horizons66.6%#13 of 32, top 41%Epoch AI

Reasoning

Grok 4 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-216%#45 of 83, top 55%Epoch AI
SimpleBench60.5%#25 of 77, top 33%Epoch AI
Kagi LLM Benchmark73.6%#16 of 99, top 17%Kagi LLM Benchmark
ARC-AGI-166.7%#44 of 83, top 54%Epoch AI
Chess Puzzles28%#38 of 129, top 30%Epoch AI2026-01-30
LMArena Hard Prompts1409#119 of 297, top 41%LMArena2026-10-08
Epoch Capabilities Index146.44#73 of 213, top 35%Epoch AI2025-07-09
ForecastBench60.9#21 of 72, top 30%Epoch AI

Math

Grok 4 Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202584%#74 of 173, top 43%Epoch AI
Omni-MATH60.3%#9 of 57, top 16%HELM Capabilities
LMArena Math1422#98 of 285, top 35%LMArena2026-10-08
FrontierMath (Feb 2025 set)19.7%#31 of 68, top 46%Epoch AI2025-11-13
FrontierMath Tier 4 (v1)2.1%#41 of 55, top 75%Epoch AI2025-08-11

Knowledge

Grok 4 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond87%#56 of 186, top 31%Epoch AI
MMLU-Pro85.1%#7 of 58, top 13%HELM Capabilities
Confabulations (lower is better)12.4%#7 of 51, top 14%Lech Mazur benchmarks
GPQA (HELM)72.7%#7 of 57, top 13%HELM Capabilities
LMArena Expert1415#111 of 273, top 41%LMArena2026-10-08

Multimodal

Grok 4 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1210#76 of 122, top 63%LMArena2026-10-09
GeoBench45%#22 of 25, top 88%Epoch AI

Multilingual

Grok 4 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1403#103 of 297, top 35%LMArena2026-10-08
LMArena Chinese1427#120 of 285, top 43%LMArena2026-10-08
LMArena French1418#103 of 223, top 47%LMArena2026-10-08
LMArena German1429#67 of 231, top 30%LMArena2026-10-08
LMArena Japanese1394#65 of 211, top 31%LMArena2026-10-08
LMArena Korean1377#82 of 213, top 39%LMArena2026-10-08
LMArena Russian1410#95 of 283, top 34%LMArena2026-10-08
LMArena Spanish1420#92 of 226, top 41%LMArena2026-10-08

Instruction Following

Grok 4 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval94.9%#2 of 57, top 4%HELM Capabilities
LMArena Instruction Following1387#116 of 298, top 39%LMArena2026-10-08

Long Context

Grok 4 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench94.4%#3 of 47, top 7%Epoch AI
LMArena Longer Query1409#109 of 291, top 38%LMArena2026-10-08

Writing & Preference

Grok 4 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1411#108 of 297, top 37%LMArena2026-10-08
LMArena Creative Writing1397#85 of 295, top 29%LMArena2026-10-08
Short-Story Creative Writing76.9%#20 of 39, top 52%Epoch AI
WildBench79.7%#32 of 57, top 57%HELM Capabilities
LMArena Multi-Turn1416#99 of 295, top 34%LMArena2026-10-08

Compare Grok 4

Other xAI models

Frequently asked questions

How good is Grok 4?

Grok 4 by xAI ranks 56th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 4th.

Is Grok 4 open source?

No. Grok 4 is proprietary and available only through xAI's API and partner platforms.

How fast is Grok 4?

Grok 4 generated about 1 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Grok 4's strengths and weaknesses?

Relative to other ranked models, Grok 4 places best in long context, instruction following, coding and lowest in multimodal, agentic & tool use, writing & preference.

What is Grok 4 best at?

Its best category is long context, where it ranks 4th on Noometry.