Model comparison

GPT-5.6 Luna vs GPT-6.1 Sol

GPT-6.1 Sol is the stronger model overall, scoring 65.6 to 54.6 on the Noometry Index. GPT-5.6 Luna costs 8.9× less per token, which makes it the better buy when GPT-6.1 Sol's lead doesn't matter for your workload.

Last verified . 33 shared benchmarks.

GPT-5.6 Luna OpenAI

54.6

Rank #30 Confirmed

GPT-6.1 Sol OpenAI

65.6

Rank #6 Confirmed

Summary

  • They share 33 benchmarks with published results for both. GPT-5.6 Luna scores higher in 1 category and GPT-6.1 Sol in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-6.1 Sol leads 81.9 to 47.6.
  • The biggest single-benchmark swing is Mystery Game Puzzles: 21% for GPT-5.6 Luna and 80% for GPT-6.1 Sol.
  • GPT-5.6 Luna is cheaper at $0.20 / $1.20 per million input/output tokens, against $2 / $10 for GPT-6.1 Sol.

Side by side

GPT-5.6 Luna and GPT-6.1 Sol specifications
GPT-5.6 LunaGPT-6.1 Sol
ProviderOpenAIOpenAI
Noometry Index54.665.6
Released2026-07-092026-09-29
WeightsProprietaryProprietary
Context window1.05M1.05M
Max output128K128K
Input $ / M tokens$0.20$2
Output $ / M tokens$1.20$10
Results tracked5234

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-6.1 Sol leads

GPT-5.6 Luna: 54.5 (#28), GPT-6.1 Sol: 63.2 (#8)

Coding benchmarks
BenchmarkGPT-5.6 LunaGPT-6.1 Sol
DeepSWE67.2%75.2%
FrontierCode39.8%50.2%
LMArena WebDev15191755
SciCode53.6%55.8%
LMArena Coding14661487
CursorBench35.9%—
WeirdML60.9%—
ALE-Bench1,667—

Agentic & Tool Use GPT-6.1 Sol leads

GPT-5.6 Luna: 34.4 (#45), GPT-6.1 Sol: 39.6 (#26)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.6 LunaGPT-6.1 Sol
APEX-Agents43%60%
GDP.pdf22.7%32%
BALROG45.6%—
Vending-Bench 24,095—

Reasoning GPT-6.1 Sol leads

GPT-5.6 Luna: 47.6 (#43), GPT-6.1 Sol: 81.9 (#2)

Reasoning benchmarks
BenchmarkGPT-5.6 LunaGPT-6.1 Sol
ARC-AGI-259.5%94.2%
NYT Connections (extended)69.4%95.5%
ARC-AGI-188%98.5%
CritPt20.6%31.7%
Chess Puzzles40%61%
LMArena Hard Prompts14511466
Mystery Game Puzzles21%80%
Epoch Capabilities Index156.39166.09
SimpleBench46.8%—
Kagi LLM Benchmark49.1%—
EBR-Bench—54.3%
DTBench89.1%—
LMCA48.5%—
Surface Evolver Bench61.9%—

Math GPT-6.1 Sol leads

GPT-5.6 Luna: 77.7 (#14), GPT-6.1 Sol: 93.7 (#1)

Math benchmarks
BenchmarkGPT-5.6 LunaGPT-6.1 Sol
FrontierMath (Tiers 1-3)82.1%93.7%
FrontierMath Tier 461%100%
OTIS Mock AIME 2024-202598.3%100%
ProofBench60%99%
LMArena Math14581464

Knowledge GPT-6.1 Sol leads

GPT-5.6 Luna: 58.5 (#34), GPT-6.1 Sol: 71.8 (#4)

Knowledge benchmarks
BenchmarkGPT-5.6 LunaGPT-6.1 Sol
GPQA Diamond91.6%95.4%
SimpleQA Verified41%73.9%
LMArena Expert14781502

Multimodal GPT-6.1 Sol leads

GPT-5.6 Luna: 42.7 (#28), GPT-6.1 Sol: 52.7 (#5)

Multimodal benchmarks
BenchmarkGPT-5.6 LunaGPT-6.1 Sol
LMArena Vision12581288
Furniture Assembly42.5%80%
Blueprint-Bench 222.6%—
LMArena Document1457—

Multilingual GPT-6.1 Sol leads

GPT-5.6 Luna: 52.8 (#78), GPT-6.1 Sol: 54.3 (#46)

Multilingual benchmarks
BenchmarkGPT-5.6 LunaGPT-6.1 Sol
LMArena Non-English14171438
LMArena Chinese14701477
LMArena Russian14281455
LMArena French1456—
LMArena German1454—
LMArena Japanese1411—
LMArena Korean1415—
LMArena Spanish1448—

Instruction Following GPT-6.1 Sol leads

GPT-5.6 Luna: 75.6 (#57), GPT-6.1 Sol: 77.0 (#29)

Instruction Following benchmarks
BenchmarkGPT-5.6 LunaGPT-6.1 Sol
LMArena Instruction Following14371468

Long Context Too close to call

GPT-5.6 Luna: 43.9 (#82), GPT-6.1 Sol: 44.9 (#54)

Long Context benchmarks
BenchmarkGPT-5.6 LunaGPT-6.1 Sol
LMArena Longer Query14361465

Writing & Preference GPT-5.6 Luna leads

GPT-5.6 Luna: 68.0 (#29), GPT-6.1 Sol: 63.6 (#63)

Writing & Preference benchmarks
BenchmarkGPT-5.6 LunaGPT-6.1 Sol
LMArena Text14311447
LMArena Creative Writing13961432
LMArena Multi-Turn14341449
EQ-Bench Creative Writing1829—
EQ-Bench 41156—

Frequently asked questions

Is GPT-5.6 Luna better than GPT-6.1 Sol?

GPT-6.1 Sol is the stronger model overall, scoring 65.6 to 54.6 on the Noometry Index. GPT-5.6 Luna costs 8.9× less per token, which makes it the better buy when GPT-6.1 Sol's lead doesn't matter for your workload.

Which is cheaper, GPT-5.6 Luna or GPT-6.1 Sol?

GPT-5.6 Luna is cheaper. It lists at $0.20 per million input tokens and $1.20 per million output tokens; GPT-6.1 Sol lists at $2 and $10.

Is GPT-5.6 Luna or GPT-6.1 Sol better for coding?

GPT-6.1 Sol scores higher on coding benchmarks: 63.2 versus 54.5 in the Noometry coding category.

Which has the bigger context window?

Both accept 1.05M tokens.

How many benchmarks do GPT-5.6 Luna and GPT-6.1 Sol share?

33 benchmarks have published results for both models. GPT-5.6 Luna has 52 scored results on Noometry and GPT-6.1 Sol has 34.

Related comparisons

Go deeper