Model comparison

Claude Opus 4 vs DeepSeek V4 Flash

DeepSeek V4 Flash is the stronger model overall, scoring 53.6 to 43.1 on the Noometry Index.

Last verified . 29 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

DeepSeek V4 Flash DeepSeek

53.6

Rank #35 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Claude Opus 4 scores higher in 1 category and DeepSeek V4 Flash in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4 Flash leads 53.7 to 27.3.
  • The biggest single-benchmark swing is ARC-AGI-1: 35.7% for Claude Opus 4 and 89% for DeepSeek V4 Flash.
  • DeepSeek V4 Flash is cheaper at $0.15 / $0.60 per million input/output tokens, against $15 / $75 for Claude Opus 4.
  • DeepSeek V4 Flash accepts more context: 1M tokens versus 200K.
  • DeepSeek V4 Flash has downloadable open weights; the other is API-only.

Side by side

Claude Opus 4 and DeepSeek V4 Flash specifications
Claude Opus 4DeepSeek V4 Flash
ProviderAnthropicDeepSeek
Noometry Index43.153.6
Released2025-05-222026-04-24
WeightsProprietaryOpen
Context window200K1M
Max output32K393K
Input $ / M tokens$15$0.15
Output $ / M tokens$75$0.60
Results tracked5641

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude Opus 4: 47.2 (#62), DeepSeek V4 Flash: 47.9 (#59)

Coding benchmarks
BenchmarkClaude Opus 4DeepSeek V4 Flash
WeirdML43.7%63%
LMArena Coding14421457
SWE-bench Verified70.7%—
FrontierCode—18.8%
SWE-bench Verified (bash only)67.6%—
Aider Polyglot72%—
LMArena WebDev—1582
SciCode—49.9%
GSO6.9%—
ALE-Bench—1,306
AlgoTune1.33—

Agentic & Tool Use Not comparable

Claude Opus 4: 34.8 (#42), DeepSeek V4 Flash: —

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4DeepSeek V4 Flash
Cybench38%—
DeepResearch Bench46.8%—
LMArena Search1127—
METR Time Horizons63.9%—

Reasoning DeepSeek V4 Flash leads

Claude Opus 4: 27.3 (#121), DeepSeek V4 Flash: 53.7 (#30)

Reasoning benchmarks
BenchmarkClaude Opus 4DeepSeek V4 Flash
ARC-AGI-28.6%61.4%
SimpleBench58.8%61.1%
Kagi LLM Benchmark74.3%52.2%
ARC-AGI-135.7%89%
CritPt0.3%16.6%
LMArena Hard Prompts13991444
DTBench81.6%90.9%
LMCA37.4%41.7%
Epoch Capabilities Index142.67154.49
NYT Connections (extended)—89.6%
Chess Puzzles—33%
EnigmaEval5.6%—
Mystery Game Puzzles—34%
ForecastBench61.1—

Math DeepSeek V4 Flash leads

Claude Opus 4: 42.0 (#86), DeepSeek V4 Flash: 60.3 (#37)

Knowledge DeepSeek V4 Flash leads

Claude Opus 4: 44.0 (#88), DeepSeek V4 Flash: 55.4 (#48)

Knowledge benchmarks
BenchmarkClaude Opus 4DeepSeek V4 Flash
GPQA Diamond76.3%91%
LMArena Expert13861441
Humanity's Last Exam10.7%—
SimpleQA Verified—33.6%
MMLU-Pro87.5%—
Confabulations15.9%—
Vectara Hallucination Rate12%—
GPQA (HELM)70.8%—

Multimodal Not comparable

Claude Opus 4: 31.5 (#106), DeepSeek V4 Flash: —

Multimodal benchmarks
BenchmarkClaude Opus 4DeepSeek V4 Flash
LMArena Vision1192—
GeoBench49%—
VPCT38%—

Multilingual DeepSeek V4 Flash leads

Claude Opus 4: 48.8 (#138), DeepSeek V4 Flash: 53.0 (#72)

Multilingual benchmarks
BenchmarkClaude Opus 4DeepSeek V4 Flash
LMArena Non-English13621420
LMArena Chinese13861468
LMArena French13721439
LMArena German13911418
LMArena Japanese13311406
LMArena Korean13211384
LMArena Russian13921428
LMArena Spanish13891436

Instruction Following Claude Opus 4 leads

Claude Opus 4: 77.1 (#28), DeepSeek V4 Flash: 74.9 (#81)

Instruction Following benchmarks
BenchmarkClaude Opus 4DeepSeek V4 Flash
LMArena Instruction Following14061421
IFEval91.8%—

Long Context DeepSeek V4 Flash leads

Claude Opus 4: 39.6 (#172), DeepSeek V4 Flash: 43.8 (#85)

Long Context benchmarks
BenchmarkClaude Opus 4DeepSeek V4 Flash
LMArena Longer Query14221434
Fiction.LiveBench61.1%—

Writing & Preference DeepSeek V4 Flash leads

Claude Opus 4: 61.2 (#89), DeepSeek V4 Flash: 63.8 (#61)

Writing & Preference benchmarks
BenchmarkClaude Opus 4DeepSeek V4 Flash
LMArena Text13771432
LMArena Creative Writing13871403
EQ-Bench Creative Writing15801559
LMArena Multi-Turn13961449
Short-Story Creative Writing83.6%—
WildBench85.2%—

Frequently asked questions

Is Claude Opus 4 better than DeepSeek V4 Flash?

DeepSeek V4 Flash is the stronger model overall, scoring 53.6 to 43.1 on the Noometry Index.

Which is cheaper, Claude Opus 4 or DeepSeek V4 Flash?

DeepSeek V4 Flash is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Claude Opus 4 lists at $15 and $75.

Is Claude Opus 4 or DeepSeek V4 Flash better for coding?

They score almost the same on coding (47.2 vs 47.9); test both on your own repository before choosing.

Which has the bigger context window?

DeepSeek V4 Flash does, with 1M tokens against 200K.

How many benchmarks do Claude Opus 4 and DeepSeek V4 Flash share?

29 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and DeepSeek V4 Flash has 41.

Related comparisons

Go deeper