Model comparison

DeepSeek V4.1 Flash vs GPT-5.6 Luna

GPT-5.6 Luna is the stronger model overall, scoring 54.6 to 52.8 on the Noometry Index. DeepSeek V4.1 Flash costs 1.7× less per token, which makes it the better buy when GPT-5.6 Luna's lead doesn't matter for your workload.

Last verified . 37 shared benchmarks.

DeepSeek V4.1 Flash DeepSeek

52.8

Rank #38 Confirmed

GPT-5.6 Luna OpenAI

54.6

Rank #30 Confirmed

Summary

  • They share 37 benchmarks with published results for both. DeepSeek V4.1 Flash scores higher in 4 categories and GPT-5.6 Luna in 6 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.6 Luna leads 77.7 to 66.7.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 26.8% for DeepSeek V4.1 Flash and 61% for GPT-5.6 Luna.
  • DeepSeek V4.1 Flash is cheaper at $0.15 / $0.60 per million input/output tokens, against $0.20 / $1.20 for GPT-5.6 Luna.
  • GPT-5.6 Luna accepts more context: 1.05M tokens versus 1M.
  • DeepSeek V4.1 Flash has downloadable open weights; the other is API-only.

Side by side

DeepSeek V4.1 Flash and GPT-5.6 Luna specifications
DeepSeek V4.1 FlashGPT-5.6 Luna
ProviderDeepSeekOpenAI
Noometry Index52.854.6
Released2026-09-092026-07-09
WeightsOpenProprietary
Context window1M1.05M
Max output393K128K
Input $ / M tokens$0.15$0.20
Output $ / M tokens$0.60$1.20
Results tracked3752

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.6 Luna leads

DeepSeek V4.1 Flash: 52.9 (#32), GPT-5.6 Luna: 54.5 (#28)

Coding benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-5.6 Luna
LMArena WebDev16191519
SciCode51.9%53.6%
LMArena Coding15061466
ALE-Bench1,0921,667
DeepSWE—67.2%
FrontierCode—39.8%
CursorBench—35.9%
WeirdML—60.9%

Agentic & Tool Use GPT-5.6 Luna leads

DeepSeek V4.1 Flash: 31.2 (#69), GPT-5.6 Luna: 34.4 (#45)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-5.6 Luna
APEX-Agents39.5%43%
GDP.pdf19.8%22.7%
BALROG—45.6%
Vending-Bench 2—4,095

Reasoning DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 50.2 (#36), GPT-5.6 Luna: 47.6 (#43)

Reasoning benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-5.6 Luna
NYT Connections (extended)89.6%69.4%
CritPt14.3%20.6%
LMArena Hard Prompts14831451
Mystery Game Puzzles43%21%
DTBench89.9%89.1%
LMCA47%48.5%
Surface Evolver Bench46.3%61.9%
Epoch Capabilities Index154.9156.39
ARC-AGI-2—59.5%
SimpleBench—46.8%
Kagi LLM Benchmark—49.1%
ARC-AGI-1—88%
Chess Puzzles—40%

Math GPT-5.6 Luna leads

DeepSeek V4.1 Flash: 66.7 (#25), GPT-5.6 Luna: 77.7 (#14)

Math benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-5.6 Luna
FrontierMath (Tiers 1-3)67.4%82.1%
FrontierMath Tier 426.8%61%
OTIS Mock AIME 2024-202598.3%98.3%
ProofBench54%60%
LMArena Math14771458

Knowledge Too close to call

DeepSeek V4.1 Flash: 57.9 (#38), GPT-5.6 Luna: 58.5 (#34)

Knowledge benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-5.6 Luna
GPQA Diamond89.8%91.6%
LMArena Expert15061478
SimpleQA Verified—41%

Multimodal GPT-5.6 Luna leads

DeepSeek V4.1 Flash: 39.1 (#61), GPT-5.6 Luna: 42.7 (#28)

Multimodal benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-5.6 Luna
LMArena Vision12771258
Furniture Assembly34.2%42.5%
Blueprint-Bench 2—22.6%
LMArena Document—1457

Multilingual DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 55.0 (#35), GPT-5.6 Luna: 52.8 (#78)

Multilingual benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-5.6 Luna
LMArena Non-English14481417
LMArena Chinese14971470
LMArena French14521456
LMArena German14841454
LMArena Japanese14121411
LMArena Korean14521415
LMArena Russian14711428
LMArena Spanish14591448

Instruction Following DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 77.3 (#26), GPT-5.6 Luna: 75.6 (#57)

Instruction Following benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-5.6 Luna
LMArena Instruction Following14741437

Long Context DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 45.2 (#47), GPT-5.6 Luna: 43.9 (#82)

Long Context benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-5.6 Luna
LMArena Longer Query14751436

Writing & Preference GPT-5.6 Luna leads

DeepSeek V4.1 Flash: 65.4 (#48), GPT-5.6 Luna: 68.0 (#29)

Writing & Preference benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-5.6 Luna
LMArena Text14621431
LMArena Creative Writing14351396
EQ-Bench Creative Writing15401829
LMArena Multi-Turn14571434
EQ-Bench 4—1156

Frequently asked questions

Is DeepSeek V4.1 Flash better than GPT-5.6 Luna?

GPT-5.6 Luna is the stronger model overall, scoring 54.6 to 52.8 on the Noometry Index. DeepSeek V4.1 Flash costs 1.7× less per token, which makes it the better buy when GPT-5.6 Luna's lead doesn't matter for your workload.

Which is cheaper, DeepSeek V4.1 Flash or GPT-5.6 Luna?

DeepSeek V4.1 Flash is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; GPT-5.6 Luna lists at $0.20 and $1.20.

Is DeepSeek V4.1 Flash or GPT-5.6 Luna better for coding?

GPT-5.6 Luna scores higher on coding benchmarks: 54.5 versus 52.9 in the Noometry coding category.

Which has the bigger context window?

GPT-5.6 Luna does, with 1.05M tokens against 1M.

How many benchmarks do DeepSeek V4.1 Flash and GPT-5.6 Luna share?

37 benchmarks have published results for both models. DeepSeek V4.1 Flash has 37 scored results on Noometry and GPT-5.6 Luna has 52.

Related comparisons

Go deeper