Model comparison

DeepSeek-V3 vs Qwen3.7 Flash

DeepSeek-V3 and Qwen3.7 Flash score almost the same on the Noometry Index (39.5 vs 39.9), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

Qwen3.7 Flash Alibaba (Qwen)

39.9

Rank #156 Confirmed

Summary

  • They share 3 benchmarks with published results for both. DeepSeek-V3 scores higher in 0 categories and Qwen3.7 Flash in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.7 Flash leads 48.9 to 37.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 37.8% for DeepSeek-V3 and 86.7% for Qwen3.7 Flash.
  • Qwen3.7 Flash is cheaper at $0.03 / $0.13 per million input/output tokens, against $0.24 / $0.90 for DeepSeek-V3.
  • Qwen3.7 Flash accepts more context: 1M tokens versus 164K.
  • DeepSeek-V3 has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3 and Qwen3.7 Flash specifications
DeepSeek-V3Qwen3.7 Flash
ProviderDeepSeekAlibaba (Qwen)
Noometry Index39.539.9
Released2024-12-262026-07-15
WeightsOpenProprietary
Context window164K1M
Max output164K131K
Input $ / M tokens$0.24$0.03
Output $ / M tokens$0.90$0.13
Results tracked607

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

DeepSeek-V3: 42.3 (#106), Qwen3.7 Flash: —

Coding benchmarks
BenchmarkDeepSeek-V3Qwen3.7 Flash
Aider Polyglot55.1%—
SciCode35.8%—
WeirdML36.1%—
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
LMArena Coding1368—
BigCodeBench Complete62.2%—
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, Qwen3.7 Flash: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3Qwen3.7 Flash
METR Time Horizons49.6%—

Reasoning Qwen3.7 Flash leads

DeepSeek-V3: 20.5 (#236), Qwen3.7 Flash: 28.2 (#108)

Reasoning benchmarks
BenchmarkDeepSeek-V3Qwen3.7 Flash
Epoch Capabilities Index135.94144.64
SimpleBench27.2%—
Kagi LLM Benchmark52.3%—
NYT Connections (extended)—43.8%
CritPt0%—
Chess Puzzles—23%
LiveBench Reasoning65.8%—
LMArena Hard Prompts1365—
Mystery Game Puzzles—15%
DTBench64.8%—
LiveBench Data Analysis60.9%—
LMCA15.5%—
BIG-Bench Hard87.5%—
ForecastBench59.1—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math Qwen3.7 Flash leads

DeepSeek-V3: 32.1 (#219), Qwen3.7 Flash: 38.3 (#140)

Math benchmarks
BenchmarkDeepSeek-V3Qwen3.7 Flash
OTIS Mock AIME 2024-202537.8%86.7%
FrontierMath (Tiers 1-3)—19.3%
Omni-MATH40.3%—
LiveBench Math73.5%—
LMArena Math1373—
MATH Level 575.5%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Qwen3.7 Flash leads

DeepSeek-V3: 37.5 (#155), Qwen3.7 Flash: 48.9 (#75)

Knowledge benchmarks
BenchmarkDeepSeek-V3Qwen3.7 Flash
GPQA Diamond67.6%82.3%
MMLU-Pro72.3%—
Confabulations26.1%—
Vectara Hallucination Rate6.1%—
GPQA (HELM)53.8%—
LMArena Expert1351—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multilingual Not comparable

DeepSeek-V3: 48.5 (#143), Qwen3.7 Flash: —

Multilingual benchmarks
BenchmarkDeepSeek-V3Qwen3.7 Flash
LMArena Non-English1358—
LMArena Chinese1391—
LMArena French1385—
LMArena German1374—
LMArena Japanese1333—
LMArena Korean1319—
LMArena Russian1373—
LMArena Spanish1358—

Instruction Following Not comparable

DeepSeek-V3: 72.8 (#130), Qwen3.7 Flash: —

Instruction Following benchmarks
BenchmarkDeepSeek-V3Qwen3.7 Flash
LiveBench Instruction Following81.5%—
IFEval83.2%—
LMArena Instruction Following1345—

Long Context Not comparable

DeepSeek-V3: 34.0 (#253), Qwen3.7 Flash: —

Long Context benchmarks
BenchmarkDeepSeek-V3Qwen3.7 Flash
Fiction.LiveBench50%—
LMArena Longer Query1352—

Writing & Preference Not comparable

DeepSeek-V3: 57.4 (#130), Qwen3.7 Flash: —

Writing & Preference benchmarks
BenchmarkDeepSeek-V3Qwen3.7 Flash
LMArena Text1375—
LMArena Creative Writing1364—
Short-Story Creative Writing77%—
EQ-Bench Creative Writing1472—
WildBench83%—
LMArena Multi-Turn1389—
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than Qwen3.7 Flash?

DeepSeek-V3 and Qwen3.7 Flash score almost the same on the Noometry Index (39.5 vs 39.9), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3 or Qwen3.7 Flash?

Qwen3.7 Flash is cheaper. It lists at $0.03 per million input tokens and $0.13 per million output tokens; DeepSeek-V3 lists at $0.24 and $0.90.

Which has the bigger context window?

Qwen3.7 Flash does, with 1M tokens against 164K.

How many benchmarks do DeepSeek-V3 and Qwen3.7 Flash share?

3 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and Qwen3.7 Flash has 7.

Related comparisons

Go deeper