Model comparison

DeepSeek-V3.2-Exp vs Llama 3.1-405B

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 30.7 on the Noometry Index.

Last verified . 25 shared benchmarks.

DeepSeek-V3.2-Exp DeepSeek

44.3

Rank #78 Confirmed

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Summary

  • They share 25 benchmarks with published results for both. DeepSeek-V3.2-Exp scores higher in 9 categories and Llama 3.1-405B in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3.2-Exp leads 62.4 to 38.9.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 87.8% for DeepSeek-V3.2-Exp and 9.7% for Llama 3.1-405B.

Side by side

DeepSeek-V3.2-Exp and Llama 3.1-405B specifications
DeepSeek-V3.2-ExpLlama 3.1-405B
ProviderDeepSeekMeta
Noometry Index44.330.7
Released2025-09-292024-07-23
WeightsOpenOpen
Context window164K—
Max output66K—
Input $ / M tokens$0.26—
Output $ / M tokens$0.38—
Results tracked4942

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 46.5 (#65), Llama 3.1-405B: 33.1 (#262)

Coding benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 3.1-405B
WeirdML39.5%21.4%
LMArena Coding14541291
SWE-bench Verified (bash only)70%—
Aider Polyglot74.2%—
LMArena WebDev1362—
SWE-bench Multilingual59%—
SciCode38.9%—

Agentic & Tool Use DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 32.7 (#59), Llama 3.1-405B: 21.0 (#140)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 3.1-405B
TheAgentCompany42.9%7.4%
Terminal-Bench39.6%—
APEX-Agents21.3%—
Berkeley Function Calling Leaderboard56.7%—
Cybench—7.5%
Vending-Bench 21,034—

Reasoning DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 22.1 (#208), Llama 3.1-405B: 16.8 (#300)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 3.1-405B
Kagi LLM Benchmark52.2%45%
LMArena Hard Prompts14341269
DTBench87.7%61.4%
Epoch Capabilities Index146.27128.75
ARC-AGI-24%—
SimpleBench—23%
NYT Connections (extended)36.7%—
ARC-AGI-157%—
CritPt2.9%—
Chess Puzzles14%—
Thematic Generalization65%—
LMCA29.1%—
BIG-Bench Hard—82.9%
ForecastBench—59.9
HellaSwag—89.2%
PIQA—85.9%
WinoGrande—89.2%

Math DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 41.7 (#87), Llama 3.1-405B: 18.4 (#290)

Math benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 3.1-405B
OTIS Mock AIME 2024-202587.8%9.7%
LMArena Math14351281
MathArena Final-Answer Competitions57.7%—
ProofBench8%—
Omni-MATH—24.9%
MATH Level 5—49.8%
FrontierMath (Feb 2025 set)22.1%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 51.7 (#66), Llama 3.1-405B: 30.4 (#227)

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 3.1-405B
GPQA Diamond83.4%50.9%
LMArena Expert14361243
MMLU-Pro—72.3%
Confabulations—17.6%
Vectara Hallucination Rate5.3%—
GPQA (HELM)—52.2%
ARC (AI2) Challenge—95.3%
MMLU—84.5%
TriviaQA—82.7%

Multilingual DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 52.2 (#90), Llama 3.1-405B: 40.7 (#214)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 3.1-405B
LMArena Non-English14091248
LMArena Chinese14611242
LMArena French14331279
LMArena German14401252
LMArena Japanese13741208
LMArena Korean13711184
LMArena Russian14241265
LMArena Spanish14401260

Instruction Following DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 74.5 (#93), Llama 3.1-405B: 65.9 (#214)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 3.1-405B
LMArena Instruction Following14131259
IFEval—81.1%

Long Context DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 47.6 (#16), Llama 3.1-405B: 38.4 (#197)

Long Context benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 3.1-405B
LMArena Longer Query14281266
Fiction.LiveBench83.3%—
CL-bench13.2%—
CL-bench Life9.5%—

Writing & Preference DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 62.4 (#77), Llama 3.1-405B: 38.9 (#251)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-ExpLlama 3.1-405B
LMArena Text14251284
LMArena Creative Writing14031262
EQ-Bench Creative Writing1515870
LMArena Multi-Turn14271297
WildBench—78.3%

Frequently asked questions

Is DeepSeek-V3.2-Exp better than Llama 3.1-405B?

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 30.7 on the Noometry Index.

Is DeepSeek-V3.2-Exp or Llama 3.1-405B better for coding?

DeepSeek-V3.2-Exp scores higher on coding benchmarks: 46.5 versus 33.1 in the Noometry coding category.

How many benchmarks do DeepSeek-V3.2-Exp and Llama 3.1-405B share?

25 benchmarks have published results for both models. DeepSeek-V3.2-Exp has 49 scored results on Noometry and Llama 3.1-405B has 42.

Related comparisons

Go deeper