Model comparison

DeepSeek-V3 vs Qwen3 32B

DeepSeek-V3 and Qwen3 32B score almost the same on the Noometry Index (39.5 vs 39.2), so choose on price, context window or the category you care about most.

Last verified . 24 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 24 benchmarks with published results for both. DeepSeek-V3 scores higher in 5 categories and Qwen3 32B in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Qwen3 32B leads 43.8 to 34.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 37.8% for DeepSeek-V3 and 66.9% for Qwen3 32B.
  • DeepSeek-V3 is cheaper at $0.24 / $0.90 per million input/output tokens, against $0.70 / $2.80 for Qwen3 32B.
  • DeepSeek-V3 accepts more context: 164K tokens versus 131K.

Side by side

DeepSeek-V3 and Qwen3 32B specifications
DeepSeek-V3Qwen3 32B
ProviderDeepSeekAlibaba (Qwen)
Noometry Index39.539.2
Released2024-12-262025-04
WeightsOpenOpen
Context window164K131K
Max output164K16K
Input $ / M tokens$0.24$0.70
Output $ / M tokens$0.90$2.80
Results tracked6026

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkDeepSeek-V3Qwen3 32B
Aider Polyglot55.1%40%
SciCode35.8%35.4%
LMArena Coding13681358
WeirdML36.1%—
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3Qwen3 32B
Berkeley Function Calling Leaderboard—48.7%
METR Time Horizons49.6%—

Reasoning Too close to call

DeepSeek-V3: 20.5 (#236), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkDeepSeek-V3Qwen3 32B
Kagi LLM Benchmark52.3%54.9%
CritPt0%0.3%
LMArena Hard Prompts13651334
DTBench64.8%67.5%
LMCA15.5%17.3%
Epoch Capabilities Index135.94138.51
SimpleBench27.2%—
Chess Puzzles—5%
LiveBench Reasoning65.8%—
LiveBench Data Analysis60.9%—
BIG-Bench Hard87.5%—
ForecastBench59.1—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math Qwen3 32B leads

DeepSeek-V3: 32.1 (#219), Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkDeepSeek-V3Qwen3 32B
OTIS Mock AIME 2024-202537.8%66.9%
LMArena Math13731399
Omni-MATH40.3%—
LiveBench Math73.5%—
MATH Level 575.5%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Qwen3 32B leads

DeepSeek-V3: 37.5 (#155), Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkDeepSeek-V3Qwen3 32B
GPQA Diamond67.6%65.7%
Vectara Hallucination Rate6.1%5.9%
LMArena Expert13511362
MMLU-Pro72.3%—
Confabulations26.1%—
GPQA (HELM)53.8%—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multilingual DeepSeek-V3 leads

DeepSeek-V3: 48.5 (#143), Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkDeepSeek-V3Qwen3 32B
LMArena Non-English13581317
LMArena Chinese13911357
LMArena German13741341
LMArena Russian13731311
LMArena French1385—
LMArena Japanese1333—
LMArena Korean1319—
LMArena Spanish1358—

Instruction Following DeepSeek-V3 leads

DeepSeek-V3: 72.8 (#130), Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkDeepSeek-V3Qwen3 32B
LMArena Instruction Following13451305
LiveBench Instruction Following81.5%—
IFEval83.2%—

Long Context Qwen3 32B leads

DeepSeek-V3: 34.0 (#253), Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkDeepSeek-V3Qwen3 32B
Fiction.LiveBench50%74.2%
LMArena Longer Query13521327

Writing & Preference DeepSeek-V3 leads

DeepSeek-V3: 57.4 (#130), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3Qwen3 32B
LMArena Text13751340
LMArena Creative Writing13641297
LMArena Multi-Turn13891331
Short-Story Creative Writing77%—
EQ-Bench Creative Writing1472—
WildBench83%—
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than Qwen3 32B?

DeepSeek-V3 and Qwen3 32B score almost the same on the Noometry Index (39.5 vs 39.2), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3 or Qwen3 32B?

DeepSeek-V3 is cheaper. It lists at $0.24 per million input tokens and $0.90 per million output tokens; Qwen3 32B lists at $0.70 and $2.80.

Is DeepSeek-V3 or Qwen3 32B better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 37.7 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3 does, with 164K tokens against 131K.

How many benchmarks do DeepSeek-V3 and Qwen3 32B share?

24 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper