Model comparison

GPT-6 Luna vs Grok-2 (Dec 2024)

GPT-6 Luna is the stronger model overall, scoring 53.3 to 33.7 on the Noometry Index.

Last verified . 21 shared benchmarks.

GPT-6 Luna OpenAI

53.3

Rank #36 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 21 benchmarks with published results for both. GPT-6 Luna scores higher in 8 categories and Grok-2 (Dec 2024) in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-6 Luna leads 76.1 to 20.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 98.9% for GPT-6 Luna and 11.5% for Grok-2 (Dec 2024).

Side by side

GPT-6 Luna and Grok-2 (Dec 2024) specifications
GPT-6 LunaGrok-2 (Dec 2024)
ProviderOpenAIxAI
Noometry Index53.333.7
Released2026-09-222024-08-13
WeightsProprietaryProprietary
Context window1.05M—
Max output128K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.50—
Results tracked4234

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-6 Luna leads

GPT-6 Luna: 55.5 (#25), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkGPT-6 LunaGrok-2 (Dec 2024)
LMArena Coding14391287
DeepSWE66.6%—
FrontierCode42.4%—
LMArena WebDev1581—
SciCode54.6%—
WeirdML—22.2%
LiveBench Coding—46.4%
ALE-Bench1,577—

Agentic & Tool Use Not comparable

GPT-6 Luna: 33.3 (#54), Grok-2 (Dec 2024): —

Agentic & Tool Use benchmarks
BenchmarkGPT-6 LunaGrok-2 (Dec 2024)
APEX-Agents44.3%—
GDP.pdf23%—

Reasoning GPT-6 Luna leads

GPT-6 Luna: 48.2 (#41), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkGPT-6 LunaGrok-2 (Dec 2024)
LMArena Hard Prompts14111272
DTBench90.1%65.2%
Epoch Capabilities Index156.28130.48
ARC-AGI-259.3%—
SimpleBench—22.7%
NYT Connections (extended)68.7%—
ARC-AGI-186.7%—
CritPt19.4%—
Chess Puzzles31%—
LiveBench Reasoning—54.8%
Mystery Game Puzzles7%—
LiveBench Data Analysis—54.5%
LMCA44.5%—
LiveBench—54.3%

Math GPT-6 Luna leads

GPT-6 Luna: 76.1 (#15), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkGPT-6 LunaGrok-2 (Dec 2024)
OTIS Mock AIME 2024-202598.9%11.5%
LMArena Math14161283
FrontierMath (Tiers 1-3)78.9%—
FrontierMath Tier 456.1%—
ProofBench64%—
LiveBench Math—54.9%
MATH Level 5—63.5%
FrontierMath (Feb 2025 set)—0.7%

Knowledge GPT-6 Luna leads

GPT-6 Luna: 57.0 (#41), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkGPT-6 LunaGrok-2 (Dec 2024)
GPQA Diamond90.5%53.8%
LMArena Expert14441254
SimpleQA Verified41.4%—
Confabulations—20.1%

Multimodal Not comparable

GPT-6 Luna: 42.4 (#30), Grok-2 (Dec 2024): —

Multimodal benchmarks
BenchmarkGPT-6 LunaGrok-2 (Dec 2024)
LMArena Vision1217—
Blueprint-Bench 231.2%—
Furniture Assembly44.2%—

Multilingual GPT-6 Luna leads

GPT-6 Luna: 50.5 (#117), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkGPT-6 LunaGrok-2 (Dec 2024)
LMArena Non-English13861282
LMArena Chinese14331289
LMArena French14201318
LMArena German13691287
LMArena Japanese13691244
LMArena Korean13601237
LMArena Russian13941286
LMArena Spanish13931281

Instruction Following GPT-6 Luna leads

GPT-6 Luna: 74.3 (#99), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkGPT-6 LunaGrok-2 (Dec 2024)
LMArena Instruction Following14091270
LiveBench Instruction Following—69.6%

Long Context GPT-6 Luna leads

GPT-6 Luna: 43.0 (#111), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkGPT-6 LunaGrok-2 (Dec 2024)
LMArena Longer Query14091276

Writing & Preference GPT-6 Luna leads

GPT-6 Luna: 58.3 (#119), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkGPT-6 LunaGrok-2 (Dec 2024)
LMArena Text13911305
LMArena Creative Writing13631284
LMArena Multi-Turn13961290
Short-Story Creative Writing—63.6%
LiveBench Language—45.6%

Frequently asked questions

Is GPT-6 Luna better than Grok-2 (Dec 2024)?

GPT-6 Luna is the stronger model overall, scoring 53.3 to 33.7 on the Noometry Index.

Is GPT-6 Luna or Grok-2 (Dec 2024) better for coding?

GPT-6 Luna scores higher on coding benchmarks: 55.5 versus 33.3 in the Noometry coding category.

How many benchmarks do GPT-6 Luna and Grok-2 (Dec 2024) share?

21 benchmarks have published results for both models. GPT-6 Luna has 42 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper