xAI, proprietary

Grok 3

Grok 3 by xAI ranks 157th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.9. Its strongest category is instruction following, where it ranks 73rd.

Last verified

Specifications

Noometry rank
#157 of 354
Index score
39.9
Evidence
Confirmed 40 results
Provider
xAI
Released
April 9, 2025
Weights
Proprietary
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
42 tokens/s Kagi
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Grok 3 category scores
  1. Coding 41.9
  2. Agentic & Tool Use 30.5
  3. Reasoning 13.7
  4. Math 38.0
  5. Knowledge 46.2
  6. Multilingual 52.3
  7. Instruction Following 75.0
  8. Long Context 38.7
  9. Writing & Preference 55.8
Grok 3 category ranks
CategoryScoreRankResults
Coding41.9#1153
Agentic & Tool Use30.5#761
Reasoning13.7#3335
Math38.0#1454
Knowledge46.2#826
Multilingual52.3#871
Instruction Following75.0#732
Long Context38.7#1922
Writing & Preference55.8#1416

Strengths and weaknesses

Categories where Grok 3 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Grok 3: strongest categories
CategoryScorevs medianRank
Instruction Following75.0+3.8#73 of 305, top 24%
Knowledge46.2+8.8#82 of 314, top 27%
Multilingual52.3+4.9#87 of 297, top 30%

Weakest categories

Grok 3: weakest categories
CategoryScorevs medianRank
Reasoning13.7−10.0#333 of 350, top 96%
Long Context38.7−2.2#192 of 296, top 65%
Agentic & Tool Use30.5+0.2#76 of 154, top 50%

Closest competitors

The models ranked just above and below Grok 3. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Grok 3
ModelRankScoreBlended $/MSpeed
Nemotron 3 Super#15340.1$0.17—Compare
Llama 3.3 Nemotron 49b Super v1#15440.1——Compare
Nemotron 3.5 Lightning#15540.0$0.0875—Compare
Qwen3.7 Flash#15639.9$0.055—Compare
GLM-4.5V#15839.8$0.9034Compare
QwQ-32B#15939.8——Compare
Step 1o Turbo 202506#16039.7——Compare
Nova 2 Lite#16139.7$0.85—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Grok 3 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot53.3%#20 of 44, top 46%Epoch AI
WeirdML37.2%#87 of 119, top 74%Epoch AI
LMArena Coding1432#109 of 294, top 38%LMArena2026-10-08

Agentic & Tool Use

Grok 3 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
BALROG29.5%#18 of 35, top 52%Epoch AI

Reasoning

Grok 3 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-20%#79 of 83, top 96%Epoch AI
SimpleBench36.1%#56 of 77, top 73%Epoch AI
Kagi LLM Benchmark61.3%#40 of 99, top 41%Kagi LLM Benchmark
ARC-AGI-15.5%#77 of 83, top 93%Epoch AI
LMArena Hard Prompts1434#87 of 297, top 30%LMArena2026-10-08
Epoch Capabilities Index138.33#114 of 213, top 54%Epoch AI2025-04-09

Math

Grok 3 Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202555.6%#110 of 173, top 64%Epoch AI2025-04-10
Omni-MATH46.4%#20 of 57, top 36%HELM Capabilities
LMArena Math1391#135 of 285, top 48%LMArena2026-10-08
MATH Level 588.7%#17 of 79, top 22%Epoch AI2025-04-10
FrontierMath (Feb 2025 set)3.8%#52 of 68, top 77%Epoch AI2025-04-10
FrontierMath Tier 4 (v1)0%#51 of 55, top 93%Epoch AI2025-07-01

Knowledge

Grok 3 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond75.8%#93 of 186, top 50%Epoch AI2025-05-26
MMLU-Pro78.8%#19 of 58, top 33%HELM Capabilities
Confabulations (lower is better)14.2%#14 of 51, top 28%no reasoningLech Mazur benchmarks
Vectara Hallucination Rate (lower is better)5.8%#20 of 96, top 21%Vectara Hallucination Leaderboard
GPQA (HELM)65%#18 of 57, top 32%HELM Capabilities
LMArena Expert1421#106 of 273, top 39%LMArena2026-10-08

Multilingual

Grok 3 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1410#87 of 297, top 30%LMArena2026-10-08
LMArena Chinese1448#103 of 285, top 37%LMArena2026-10-08
LMArena French1460#49 of 223, top 22%LMArena2026-10-08
LMArena German1431#65 of 231, top 29%LMArena2026-10-08
LMArena Japanese1387#71 of 211, top 34%LMArena2026-10-08
LMArena Korean1373#83 of 213, top 39%LMArena2026-10-08
LMArena Russian1416#86 of 283, top 31%LMArena2026-10-08
LMArena Spanish1417#94 of 226, top 42%LMArena2026-10-08

Instruction Following

Grok 3 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval88.4%#12 of 57, top 22%HELM Capabilities
LMArena Instruction Following1409#89 of 298, top 30%LMArena2026-10-08

Long Context

Grok 3 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench58.3%#31 of 47, top 66%Epoch AI
LMArena Longer Query1439#66 of 291, top 23%LMArena2026-10-08

Writing & Preference

Grok 3 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1426#86 of 297, top 29%LMArena2026-10-08
LMArena Creative Writing1414#60 of 295, top 21%LMArena2026-10-08
Short-Story Creative Writing76.4%#22 of 39, top 57%Epoch AI
EQ-Bench Creative Writing1186#86 of 115, top 75%EQ-Bench
WildBench84.9%#13 of 57, top 23%HELM Capabilities
LMArena Multi-Turn1425#91 of 295, top 31%LMArena2026-10-08

Compare Grok 3

Other xAI models

Frequently asked questions

How good is Grok 3?

Grok 3 by xAI ranks 157th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.9. Its strongest category is instruction following, where it ranks 73rd.

Is Grok 3 open source?

No. Grok 3 is proprietary and available only through xAI's API and partner platforms.

How fast is Grok 3?

Grok 3 generated about 42 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Grok 3's strengths and weaknesses?

Relative to other ranked models, Grok 3 places best in instruction following, knowledge, multilingual and lowest in reasoning, long context, agentic & tool use.

What is Grok 3 best at?

Its best category is instruction following, where it ranks 73rd on Noometry.