Model comparison

Grok 4.20 (Non-Reasoning) vs Grok 4.3

Grok 4.20 (Non-Reasoning) is the stronger model overall, scoring 48.6 to 43.8 on the Noometry Index.

Last verified . 36 shared benchmarks.

Grok 4.20 (Non-Reasoning) xAI

48.6

Rank #54 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 36 benchmarks with published results for both. Grok 4.20 (Non-Reasoning) scores higher in 10 categories and Grok 4.3 in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.20 (Non-Reasoning) leads 52.3 to 35.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 85.4% for Grok 4.20 (Non-Reasoning) and 55.2% for Grok 4.3.
  • Both cost about the same: $1.25 input and $2.50 output per million tokens.

Side by side

Grok 4.20 (Non-Reasoning) and Grok 4.3 specifications
Grok 4.20 (Non-Reasoning)Grok 4.3
ProviderxAIxAI
Noometry Index48.643.8
Released2026-02-172026-04-17
WeightsProprietaryProprietary
Context window1M1M
Max output30K30K
Input $ / M tokens$1.25$1.25
Output $ / M tokens$2.50$2.50
Results tracked4640

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.20 (Non-Reasoning): 42.1 (#112), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Grok 4.3
LMArena WebDev13751357
WeirdML52.3%49.9%
LMArena Coding14591415
ALE-Bench1,150944.17
SciCode—47.3%

Agentic & Tool Use Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 34.4 (#46), Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Grok 4.3
LMArena Search11891165
Vending-Bench 24,66335.26
Terminal-Bench57.3%—
τ²-bench Banking18%—
GDP.pdf—8%

Reasoning Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 52.3 (#32), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Grok 4.3
NYT Connections (extended)85.4%55.2%
Chess Puzzles24%25%
LMArena Hard Prompts14511396
DTBench90.1%90.7%
LMCA38.7%38.3%
Epoch Capabilities Index151.98149.16
ForecastBench61.460.3
ARC-AGI-265.1%—
Kagi LLM Benchmark75%—
ARC-AGI-189.5%—
CritPt—8%
Thematic Generalization63.8%—

Math Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 48.2 (#65), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Grok 4.3
FrontierMath (Tiers 1-3)44.9%42.8%
FrontierMath Tier 417.1%14.6%
OTIS Mock AIME 2024-202592.2%93.3%
ProofBench14%11%
LMArena Math14551388

Knowledge Too close to call

Grok 4.20 (Non-Reasoning): 52.8 (#60), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Grok 4.3
GPQA Diamond89.3%88.8%
SimpleQA Verified30.2%33.2%
LMArena Expert14391385

Multimodal Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 33.3 (#98), Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Grok 4.3
LMArena Vision12631229
Blueprint-Bench 20%0%
LMArena Document1416—

Multilingual Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 54.5 (#40), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Grok 4.3
LMArena Non-English14411385
LMArena Chinese14811422
LMArena French14761412
LMArena German14651395
LMArena Japanese14491379
LMArena Korean14171356
LMArena Russian14581399
LMArena Spanish14431398

Instruction Following Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 74.8 (#83), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Grok 4.3
LMArena Instruction Following14201366

Long Context Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 45.5 (#34), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Grok 4.3
LMArena Longer Query14371393
CL-bench22.2%—
CL-bench Life11.9%—

Writing & Preference Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 65.7 (#44), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Grok 4.3
LMArena Text14511397
LMArena Creative Writing14381380
LMArena Multi-Turn14561406
EQ-Bench Creative Writing1574—
EQ-Bench 4—1075

Frequently asked questions

Is Grok 4.20 (Non-Reasoning) better than Grok 4.3?

Grok 4.20 (Non-Reasoning) is the stronger model overall, scoring 48.6 to 43.8 on the Noometry Index.

Which is cheaper, Grok 4.20 (Non-Reasoning) or Grok 4.3?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; Grok 4.20 (Non-Reasoning) lists at $1.25 and $2.50.

Is Grok 4.20 (Non-Reasoning) or Grok 4.3 better for coding?

They score almost the same on coding (42.1 vs 41.6); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do Grok 4.20 (Non-Reasoning) and Grok 4.3 share?

36 benchmarks have published results for both models. Grok 4.20 (Non-Reasoning) has 46 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper