Model comparison

GPT-5.6 Luna vs Grok 4.5

GPT-5.6 Luna and Grok 4.5 score almost the same on the Noometry Index (54.6 vs 55.0), so choose on price, context window or the category you care about most.

Last verified . 48 shared benchmarks.

GPT-5.6 Luna OpenAI

54.6

Rank #30 Confirmed

Grok 4.5 xAI

55.0

Rank #25 Confirmed

Summary

  • They share 48 benchmarks with published results for both. GPT-5.6 Luna scores higher in 4 categories and Grok 4.5 in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.6 Luna leads 77.7 to 60.9.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 61% for GPT-5.6 Luna and 24.4% for Grok 4.5.
  • GPT-5.6 Luna is cheaper at $0.20 / $1.20 per million input/output tokens, against $2 / $6 for Grok 4.5.
  • GPT-5.6 Luna accepts more context: 1.05M tokens versus 500K.

Side by side

GPT-5.6 Luna and Grok 4.5 specifications
GPT-5.6 LunaGrok 4.5
ProviderOpenAIxAI
Noometry Index54.655.0
Released2026-07-092026-07-08
WeightsProprietaryProprietary
Context window1.05M500K
Max output128K500K
Input $ / M tokens$0.20$2
Output $ / M tokens$1.20$6
Results tracked5252

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.6 Luna leads

GPT-5.6 Luna: 54.5 (#28), Grok 4.5: 52.2 (#35)

Coding benchmarks
BenchmarkGPT-5.6 LunaGrok 4.5
DeepSWE67.2%53.8%
FrontierCode39.8%42.4%
LMArena WebDev15191553
SciCode53.6%54.1%
WeirdML60.9%46.4%
LMArena Coding14661474
ALE-Bench1,6671,309
CursorBench35.9%—

Agentic & Tool Use Grok 4.5 leads

GPT-5.6 Luna: 34.4 (#45), Grok 4.5: 44.4 (#17)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.6 LunaGrok 4.5
APEX-Agents43%56.2%
GDP.pdf22.7%14%
Vending-Bench 24,0953,887
τ²-bench Banking—47.9%
PostTrainBench—23.4%
BALROG45.6%—
GBAEval—65.4%
LMArena Search—1213

Reasoning Grok 4.5 leads

GPT-5.6 Luna: 47.6 (#43), Grok 4.5: 56.1 (#25)

Reasoning benchmarks
BenchmarkGPT-5.6 LunaGrok 4.5
ARC-AGI-259.5%52.6%
SimpleBench46.8%70%
Kagi LLM Benchmark49.1%83.5%
NYT Connections (extended)69.4%79.9%
ARC-AGI-188%87.2%
CritPt20.6%15.4%
Chess Puzzles40%36%
LMArena Hard Prompts14511462
DTBench89.1%96.5%
LMCA48.5%45.2%
Surface Evolver Bench61.9%74.4%
Epoch Capabilities Index156.39153.92
Mystery Game Puzzles21%—

Math GPT-5.6 Luna leads

GPT-5.6 Luna: 77.7 (#14), Grok 4.5: 60.9 (#35)

Math benchmarks
BenchmarkGPT-5.6 LunaGrok 4.5
FrontierMath (Tiers 1-3)82.1%57.2%
FrontierMath Tier 461%24.4%
OTIS Mock AIME 2024-202598.3%97.8%
ProofBench60%31%
LMArena Math14581459

Knowledge Grok 4.5 leads

GPT-5.6 Luna: 58.5 (#34), Grok 4.5: 62.3 (#24)

Knowledge benchmarks
BenchmarkGPT-5.6 LunaGrok 4.5
GPQA Diamond91.6%93.4%
SimpleQA Verified41%48.3%
LMArena Expert14781466

Multimodal GPT-5.6 Luna leads

GPT-5.6 Luna: 42.7 (#28), Grok 4.5: 37.6 (#72)

Multimodal benchmarks
BenchmarkGPT-5.6 LunaGrok 4.5
LMArena Vision12581288
Blueprint-Bench 222.6%27.3%
Furniture Assembly42.5%22.5%
LMArena Document14571452

Multilingual Grok 4.5 leads

GPT-5.6 Luna: 52.8 (#78), Grok 4.5: 54.4 (#42)

Multilingual benchmarks
BenchmarkGPT-5.6 LunaGrok 4.5
LMArena Non-English14171440
LMArena Chinese14701496
LMArena French14561456
LMArena German14541446
LMArena Japanese14111428
LMArena Korean14151404
LMArena Russian14281448
LMArena Spanish14481450

Instruction Following Too close to call

GPT-5.6 Luna: 75.6 (#57), Grok 4.5: 76.0 (#48)

Instruction Following benchmarks
BenchmarkGPT-5.6 LunaGrok 4.5
LMArena Instruction Following14371446

Long Context Too close to call

GPT-5.6 Luna: 43.9 (#82), Grok 4.5: 44.8 (#56)

Long Context benchmarks
BenchmarkGPT-5.6 LunaGrok 4.5
LMArena Longer Query14361463

Writing & Preference GPT-5.6 Luna leads

GPT-5.6 Luna: 68.0 (#29), Grok 4.5: 65.8 (#42)

Writing & Preference benchmarks
BenchmarkGPT-5.6 LunaGrok 4.5
LMArena Text14311448
LMArena Creative Writing13961442
EQ-Bench Creative Writing18291579
LMArena Multi-Turn14341456
EQ-Bench 41156—

Frequently asked questions

Is GPT-5.6 Luna better than Grok 4.5?

GPT-5.6 Luna and Grok 4.5 score almost the same on the Noometry Index (54.6 vs 55.0), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.6 Luna or Grok 4.5?

GPT-5.6 Luna is cheaper. It lists at $0.20 per million input tokens and $1.20 per million output tokens; Grok 4.5 lists at $2 and $6.

Is GPT-5.6 Luna or Grok 4.5 better for coding?

GPT-5.6 Luna scores higher on coding benchmarks: 54.5 versus 52.2 in the Noometry coding category.

Which has the bigger context window?

GPT-5.6 Luna does, with 1.05M tokens against 500K.

How many benchmarks do GPT-5.6 Luna and Grok 4.5 share?

48 benchmarks have published results for both models. GPT-5.6 Luna has 52 scored results on Noometry and Grok 4.5 has 52.

Related comparisons

Go deeper