Model comparison

GPT-5.4 nano vs Grok 4.6

Grok 4.6 is the stronger model overall, scoring 56.9 to 41.9 on the Noometry Index. GPT-5.4 nano costs 6.5× less per token, which makes it the better buy when Grok 4.6's lead doesn't matter for your workload.

Last verified . 35 shared benchmarks.

GPT-5.4 nano OpenAI

41.9

Rank #125 Confirmed

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Summary

  • They share 35 benchmarks with published results for both. GPT-5.4 nano scores higher in 0 categories and Grok 4.6 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.6 leads 61.4 to 23.7.
  • The biggest single-benchmark swing is ARC-AGI-2: 5.7% for GPT-5.4 nano and 67.1% for Grok 4.6.
  • GPT-5.4 nano is cheaper at $0.20 / $1.25 per million input/output tokens, against $2 / $6 for Grok 4.6.
  • Grok 4.6 accepts more context: 500K tokens versus 400K.

Side by side

GPT-5.4 nano and Grok 4.6 specifications
GPT-5.4 nanoGrok 4.6
ProviderOpenAIxAI
Noometry Index41.956.9
Released2026-03-172026-08-12
WeightsProprietaryProprietary
Context window400K500K
Max output128K500K
Input $ / M tokens$0.20$2
Output $ / M tokens$1.25$6
Results tracked4049

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

GPT-5.4 nano: 43.6 (#84), Grok 4.6: 58.5 (#16)

Coding benchmarks
BenchmarkGPT-5.4 nanoGrok 4.6
SciCode46.9%56.5%
WeirdML49.2%67.3%
LMArena Coding14051465
ALE-Bench1,0051,508
DeepSWE—67.5%
FrontierCode—48%
CursorBench—41.4%
LMArena WebDev—1617
FrontierSWE—25.3%

Agentic & Tool Use Not comparable

GPT-5.4 nano: —, Grok 4.6: 39.4 (#27)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.4 nanoGrok 4.6
APEX-Agents—65.3%
GDP.pdf—17.2%
Vending-Bench 2—9,047

Reasoning Grok 4.6 leads

GPT-5.4 nano: 23.7 (#173), Grok 4.6: 61.4 (#20)

Reasoning benchmarks
BenchmarkGPT-5.4 nanoGrok 4.6
ARC-AGI-25.7%67.1%
ARC-AGI-151.5%87.5%
CritPt9.3%19.7%
Chess Puzzles30%40%
LMArena Hard Prompts13811447
Mystery Game Puzzles9%34%
DTBench80.3%97.3%
LMCA36.9%48.5%
Epoch Capabilities Index145.81156.44
SimpleBench—75.9%
Kagi LLM Benchmark39.7%—
NYT Connections (extended)—80%
EBR-Bench—30.5%
ForecastBench57.3—

Math Grok 4.6 leads

GPT-5.4 nano: 40.9 (#88), Grok 4.6: 67.0 (#24)

Math benchmarks
BenchmarkGPT-5.4 nanoGrok 4.6
FrontierMath (Tiers 1-3)44.9%66%
FrontierMath Tier 412.2%31.7%
OTIS Mock AIME 2024-202587.8%99.2%
ProofBench5%51%
LMArena Math14061423
FrontierMath (Feb 2025 set)25.9%—
FrontierMath Tier 4 (v1)6.3%—

Knowledge Grok 4.6 leads

GPT-5.4 nano: 41.9 (#103), Grok 4.6: 63.3 (#20)

Knowledge benchmarks
BenchmarkGPT-5.4 nanoGrok 4.6
GPQA Diamond78.5%94%
SimpleQA Verified11.7%49.3%
LMArena Expert13961467
Vectara Hallucination Rate3.1%—

Multimodal Grok 4.6 leads

GPT-5.4 nano: 36.7 (#78), Grok 4.6: 43.6 (#23)

Multimodal benchmarks
BenchmarkGPT-5.4 nanoGrok 4.6
LMArena Vision11961263
Blueprint-Bench 2—33.2%
Furniture Assembly—40%
LMArena Document—1452

Multilingual Grok 4.6 leads

GPT-5.4 nano: 48.6 (#140), Grok 4.6: 53.0 (#74)

Multilingual benchmarks
BenchmarkGPT-5.4 nanoGrok 4.6
LMArena Non-English13591420
LMArena Chinese13921480
LMArena French13961461
LMArena German13671431
LMArena Japanese13431376
LMArena Korean13201397
LMArena Russian13631422
LMArena Spanish13711404

Instruction Following Grok 4.6 leads

GPT-5.4 nano: 71.9 (#144), Grok 4.6: 75.4 (#63)

Instruction Following benchmarks
BenchmarkGPT-5.4 nanoGrok 4.6
LMArena Instruction Following13621431

Long Context Grok 4.6 leads

GPT-5.4 nano: 41.6 (#137), Grok 4.6: 44.5 (#66)

Long Context benchmarks
BenchmarkGPT-5.4 nanoGrok 4.6
LMArena Longer Query13661454

Writing & Preference Grok 4.6 leads

GPT-5.4 nano: 55.7 (#142), Grok 4.6: 62.3 (#80)

Writing & Preference benchmarks
BenchmarkGPT-5.4 nanoGrok 4.6
LMArena Text13721428
LMArena Creative Writing13141428
LMArena Multi-Turn13821425

Frequently asked questions

Is GPT-5.4 nano better than Grok 4.6?

Grok 4.6 is the stronger model overall, scoring 56.9 to 41.9 on the Noometry Index. GPT-5.4 nano costs 6.5× less per token, which makes it the better buy when Grok 4.6's lead doesn't matter for your workload.

Which is cheaper, GPT-5.4 nano or Grok 4.6?

GPT-5.4 nano is cheaper. It lists at $0.20 per million input tokens and $1.25 per million output tokens; Grok 4.6 lists at $2 and $6.

Is GPT-5.4 nano or Grok 4.6 better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 43.6 in the Noometry coding category.

Which has the bigger context window?

Grok 4.6 does, with 500K tokens against 400K.

How many benchmarks do GPT-5.4 nano and Grok 4.6 share?

35 benchmarks have published results for both models. GPT-5.4 nano has 40 scored results on Noometry and Grok 4.6 has 49.

Related comparisons

Go deeper