Model comparison

Claude Opus 4.6 vs DeepSeek V4 Flash

Claude Opus 4.6 is the stronger model overall, scoring 58.2 to 53.6 on the Noometry Index. DeepSeek V4 Flash costs 38× less per token, which makes it the better buy when Claude Opus 4.6's lead doesn't matter for your workload.

Last verified . 39 shared benchmarks.

Claude Opus 4.6 Anthropic

58.2

Rank #20 Confirmed

DeepSeek V4 Flash DeepSeek

53.6

Rank #35 Confirmed

Summary

  • They share 39 benchmarks with published results for both. Claude Opus 4.6 scores higher in 8 categories and DeepSeek V4 Flash in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Claude Opus 4.6 leads 73.5 to 63.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 83.6% for Claude Opus 4.6 and 52.2% for DeepSeek V4 Flash.
  • DeepSeek V4 Flash is cheaper at $0.15 / $0.60 per million input/output tokens, against $5 / $25 for Claude Opus 4.6.
  • DeepSeek V4 Flash has downloadable open weights; the other is API-only.

Side by side

Claude Opus 4.6 and DeepSeek V4 Flash specifications
Claude Opus 4.6DeepSeek V4 Flash
ProviderAnthropicDeepSeek
Noometry Index58.253.6
Released2026-02-042026-04-24
WeightsProprietaryOpen
Context window1M1M
Max output128K393K
Input $ / M tokens$5$0.15
Output $ / M tokens$25$0.60
Results tracked6841

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.6 leads

Claude Opus 4.6: 57.2 (#20), DeepSeek V4 Flash: 47.9 (#59)

Coding benchmarks
BenchmarkClaude Opus 4.6DeepSeek V4 Flash
FrontierCode26.6%18.8%
LMArena WebDev15471582
WeirdML78%63%
LMArena Coding15361457
ALE-Bench996.51,306
SWE-bench Verified78.7%—
SWE-bench Verified (bash only)75.6%—
SWE-bench Multilingual72%—
SciCode—49.9%
GSO41.2%—
AlgoTune1.47—

Agentic & Tool Use Not comparable

Claude Opus 4.6: 51.1 (#4), DeepSeek V4 Flash: —

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.6DeepSeek V4 Flash
Terminal-Bench79.8%—
APEX-Agents46.3%—
Remote Labor Index4.2%—
τ²-bench Banking27.3%—
Cybench93%—
DeepResearch Bench55.3%—
GBAEval44.1%—
LMArena Search1253—
METR Time Horizons78.9%—
Vending-Bench 28,018—

Reasoning Claude Opus 4.6 leads

Claude Opus 4.6: 57.8 (#23), DeepSeek V4 Flash: 53.7 (#30)

Reasoning benchmarks
BenchmarkClaude Opus 4.6DeepSeek V4 Flash
ARC-AGI-269.2%61.4%
SimpleBench67.6%61.1%
Kagi LLM Benchmark83.6%52.2%
NYT Connections (extended)92.1%89.6%
ARC-AGI-194%89%
Chess Puzzles17%33%
LMArena Hard Prompts15271444
Mystery Game Puzzles25%34%
DTBench91.2%90.9%
LMCA55.8%41.7%
Epoch Capabilities Index155.24154.49
CritPt—16.6%
EnigmaEval7.6%—
Thematic Generalization80.6%—
EBR-Bench12.7%—
ForecastBench60—

Math Claude Opus 4.6 leads

Claude Opus 4.6: 63.0 (#31), DeepSeek V4 Flash: 60.3 (#37)

Math benchmarks
BenchmarkClaude Opus 4.6DeepSeek V4 Flash
FrontierMath (Tiers 1-3)66%57.5%
FrontierMath Tier 426.8%24.4%
MathArena Final-Answer Competitions78.5%76.5%
OTIS Mock AIME 2024-202594.4%94.4%
ProofBench50%56%
LMArena Math15191427
FrontierMath (Feb 2025 set)40.7%—
FrontierMath Tier 4 (v1)22.9%—

Knowledge Claude Opus 4.6 leads

Claude Opus 4.6: 61.9 (#26), DeepSeek V4 Flash: 55.4 (#48)

Knowledge benchmarks
BenchmarkClaude Opus 4.6DeepSeek V4 Flash
GPQA Diamond90.5%91%
SimpleQA Verified47%33.6%
LMArena Expert15461441
Humanity's Last Exam34.4%—
Vectara Hallucination Rate12.2%—

Multimodal Not comparable

Claude Opus 4.6: 37.3 (#74), DeepSeek V4 Flash: —

Multimodal benchmarks
BenchmarkClaude Opus 4.6DeepSeek V4 Flash
LMArena Vision1316—
Furniture Assembly28.3%—
LMArena Document1507—

Multilingual Claude Opus 4.6 leads

Claude Opus 4.6: 57.9 (#6), DeepSeek V4 Flash: 53.0 (#72)

Multilingual benchmarks
BenchmarkClaude Opus 4.6DeepSeek V4 Flash
LMArena Non-English14891420
LMArena Chinese15511468
LMArena French15131439
LMArena German15021418
LMArena Japanese14841406
LMArena Korean14641384
LMArena Russian14971428
LMArena Spanish15101436

Instruction Following Claude Opus 4.6 leads

Claude Opus 4.6: 79.5 (#4), DeepSeek V4 Flash: 74.9 (#81)

Instruction Following benchmarks
BenchmarkClaude Opus 4.6DeepSeek V4 Flash
LMArena Instruction Following15231421

Long Context Claude Opus 4.6 leads

Claude Opus 4.6: 48.1 (#13), DeepSeek V4 Flash: 43.8 (#85)

Long Context benchmarks
BenchmarkClaude Opus 4.6DeepSeek V4 Flash
LMArena Longer Query15201434
CL-bench20.7%—
CL-bench Life17%—

Writing & Preference Claude Opus 4.6 leads

Claude Opus 4.6: 73.5 (#10), DeepSeek V4 Flash: 63.8 (#61)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.6DeepSeek V4 Flash
LMArena Text15031432
LMArena Creative Writing15051403
EQ-Bench Creative Writing18091559
LMArena Multi-Turn15131449
EQ-Bench 41223—

Frequently asked questions

Is Claude Opus 4.6 better than DeepSeek V4 Flash?

Claude Opus 4.6 is the stronger model overall, scoring 58.2 to 53.6 on the Noometry Index. DeepSeek V4 Flash costs 38× less per token, which makes it the better buy when Claude Opus 4.6's lead doesn't matter for your workload.

Which is cheaper, Claude Opus 4.6 or DeepSeek V4 Flash?

DeepSeek V4 Flash is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Claude Opus 4.6 lists at $5 and $25.

Is Claude Opus 4.6 or DeepSeek V4 Flash better for coding?

Claude Opus 4.6 scores higher on coding benchmarks: 57.2 versus 47.9 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do Claude Opus 4.6 and DeepSeek V4 Flash share?

39 benchmarks have published results for both models. Claude Opus 4.6 has 68 scored results on Noometry and DeepSeek V4 Flash has 41.

Related comparisons

Go deeper