Model comparison

DeepSeek-V3.2-Exp vs Llama 4 Maverick

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 30.9 on the Noometry Index.

Last verified . 36 shared benchmarks.

DeepSeek-V3.2-Exp DeepSeek

44.3

Rank #78 Confirmed

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Summary

  • They share 36 benchmarks with published results for both. DeepSeek-V3.2-Exp scores higher in 9 categories and Llama 4 Maverick in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3.2-Exp leads 62.4 to 38.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 87.8% for DeepSeek-V3.2-Exp and 20.6% for Llama 4 Maverick.
  • Both cost about the same: $0.26 input and $0.38 output per million tokens.
  • DeepSeek-V3.2-Exp accepts more context: 164K tokens versus 128K.

Side by side

DeepSeek-V3.2-Exp and Llama 4 Maverick specifications
DeepSeek-V3.2-ExpLlama 4 Maverick
ProviderDeepSeekMeta
Noometry Index44.330.9
Released2025-09-292025-04-05
WeightsOpenOpen
Context window164K128K
Max output66K4K
Input $ / M tokens$0.26$0.19
Output $ / M tokens$0.38$0.65
Results tracked4954

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 46.5 (#65), Llama 4 Maverick: 26.6 (#324)

Coding benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 4 Maverick
SWE-bench Verified (bash only)70%21%
Aider Polyglot74.2%15.6%
SciCode38.9%33.1%
WeirdML39.5%24.5%
LMArena Coding14541302
LMArena WebDev1362—
SWE-bench Multilingual59%—
BigCodeBench Instruct—49.7%
BigCodeBench Complete—61.4%
ALE-Bench—172.97

Agentic & Tool Use DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 32.7 (#59), Llama 4 Maverick: 28.2 (#91)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 4 Maverick
Berkeley Function Calling Leaderboard56.7%37.3%
Terminal-Bench39.6%—
APEX-Agents21.3%—
TheAgentCompany42.9%—
Vending-Bench 21,034—

Reasoning DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 22.1 (#208), Llama 4 Maverick: 10.1 (#342)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 4 Maverick
ARC-AGI-24%0%
Kagi LLM Benchmark52.2%55.9%
NYT Connections (extended)36.7%8%
ARC-AGI-157%4.4%
CritPt2.9%0%
LMArena Hard Prompts14341281
DTBench87.7%61.9%
LMCA29.1%15.9%
Epoch Capabilities Index146.27132.2
SimpleBench—27.7%
Chess Puzzles14%—
EnigmaEval—0.6%
Thematic Generalization65%—
ForecastBench—57.5

Math DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 41.7 (#87), Llama 4 Maverick: 26.0 (#262)

Math benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 4 Maverick
OTIS Mock AIME 2024-202587.8%20.6%
LMArena Math14351299
FrontierMath (Feb 2025 set)22.1%0.7%
MathArena Final-Answer Competitions57.7%—
ProofBench8%—
Omni-MATH—42.2%
MATH Level 5—73%
FrontierMath Tier 4 (v1)2.1%—

Knowledge DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 51.7 (#66), Llama 4 Maverick: 33.4 (#204)

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 4 Maverick
GPQA Diamond83.4%67%
Vectara Hallucination Rate5.3%8.2%
LMArena Expert14361259
Humanity's Last Exam—5.7%
MMLU-Pro—81%
Confabulations—22.6%
GPQA (HELM)—65%

Multimodal Not comparable

DeepSeek-V3.2-Exp: —, Llama 4 Maverick: 31.6 (#105)

Multimodal benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 4 Maverick
LMArena Vision—1142
GeoBench—52%
SpatialViz-Bench—31.8%

Multilingual DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 52.2 (#90), Llama 4 Maverick: 42.2 (#195)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 4 Maverick
LMArena Non-English14091269
LMArena Chinese14611277
LMArena French14331259
LMArena German14401291
LMArena Japanese13741207
LMArena Korean13711203
LMArena Russian14241286
LMArena Spanish14401293

Instruction Following DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 74.5 (#93), Llama 4 Maverick: 71.7 (#146)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 4 Maverick
LMArena Instruction Following14131267
IFEval—90.8%

Long Context DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 47.6 (#16), Llama 4 Maverick: 31.4 (#279)

Long Context benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 4 Maverick
Fiction.LiveBench83.3%46.2%
LMArena Longer Query14281280
CL-bench13.2%—
CL-bench Life9.5%—

Writing & Preference DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 62.4 (#77), Llama 4 Maverick: 38.8 (#252)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 4 Maverick
LMArena Text14251287
LMArena Creative Writing14031267
EQ-Bench Creative Writing1515860
LMArena Multi-Turn14271289
Short-Story Creative Writing—62%
WildBench—80%

Frequently asked questions

Is DeepSeek-V3.2-Exp better than Llama 4 Maverick?

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 30.9 on the Noometry Index.

Which is cheaper, DeepSeek-V3.2-Exp or Llama 4 Maverick?

DeepSeek-V3.2-Exp is cheaper. It lists at $0.26 per million input tokens and $0.38 per million output tokens; Llama 4 Maverick lists at $0.19 and $0.65.

Is DeepSeek-V3.2-Exp or Llama 4 Maverick better for coding?

DeepSeek-V3.2-Exp scores higher on coding benchmarks: 46.5 versus 26.6 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3.2-Exp does, with 164K tokens against 128K.

How many benchmarks do DeepSeek-V3.2-Exp and Llama 4 Maverick share?

36 benchmarks have published results for both models. DeepSeek-V3.2-Exp has 49 scored results on Noometry and Llama 4 Maverick has 54.

Related comparisons

Go deeper