Model comparison

DeepSeek-V3.1 vs Llama 4 Maverick

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 30.9 on the Noometry Index.

Last verified . 27 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Summary

  • They share 27 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 8 categories and Llama 4 Maverick in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3.1 leads 60.3 to 38.8.
  • The biggest single-benchmark swing is DTBench: 82.7% for DeepSeek-V3.1 and 61.9% for Llama 4 Maverick.
  • Llama 4 Maverick is cheaper at $0.19 / $0.65 per million input/output tokens, against $0.25 / $0.95 for DeepSeek-V3.1.
  • DeepSeek-V3.1 accepts more context: 164K tokens versus 128K.

Side by side

DeepSeek-V3.1 and Llama 4 Maverick specifications
DeepSeek-V3.1Llama 4 Maverick
ProviderDeepSeekMeta
Noometry Index42.830.9
Released2025-08-212025-04-05
WeightsOpenOpen
Context window164K128K
Max output8K4K
Input $ / M tokens$0.25$0.19
Output $ / M tokens$0.95$0.65
Results tracked2754

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.1 leads

DeepSeek-V3.1: 40.3 (#144), Llama 4 Maverick: 26.6 (#324)

Coding benchmarks
BenchmarkDeepSeek-V3.1Llama 4 Maverick
WeirdML38.4%24.5%
LMArena Coding14171302
SWE-bench Verified (bash only)—21%
Aider Polyglot—15.6%
SciCode—33.1%
BigCodeBench Instruct—49.7%
BigCodeBench Complete—61.4%
ALE-Bench—172.97

Agentic & Tool Use Not comparable

DeepSeek-V3.1: —, Llama 4 Maverick: 28.2 (#91)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.1Llama 4 Maverick
Berkeley Function Calling Leaderboard—37.3%

Reasoning DeepSeek-V3.1 leads

DeepSeek-V3.1: 27.9 (#110), Llama 4 Maverick: 10.1 (#342)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1Llama 4 Maverick
SimpleBench40%27.7%
Kagi LLM Benchmark53.2%55.9%
LMArena Hard Prompts14171281
DTBench82.7%61.9%
LMCA24.3%15.9%
Epoch Capabilities Index139.92132.2
ForecastBench5857.5
ARC-AGI-2—0%
NYT Connections (extended)—8%
ARC-AGI-1—4.4%
CritPt—0%
EnigmaEval—0.6%

Math DeepSeek-V3.1 leads

DeepSeek-V3.1: 38.9 (#122), Llama 4 Maverick: 26.0 (#262)

Math benchmarks
BenchmarkDeepSeek-V3.1Llama 4 Maverick
LMArena Math14201299
OTIS Mock AIME 2024-2025—20.6%
Omni-MATH—42.2%
MATH Level 5—73%
FrontierMath (Feb 2025 set)—0.7%

Knowledge DeepSeek-V3.1 leads

DeepSeek-V3.1: 43.7 (#90), Llama 4 Maverick: 33.4 (#204)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1Llama 4 Maverick
Vectara Hallucination Rate5.5%8.2%
LMArena Expert14051259
GPQA Diamond—67%
Humanity's Last Exam—5.7%
MMLU-Pro—81%
Confabulations—22.6%
GPQA (HELM)—65%

Multimodal Not comparable

DeepSeek-V3.1: —, Llama 4 Maverick: 31.6 (#105)

Multimodal benchmarks
BenchmarkDeepSeek-V3.1Llama 4 Maverick
LMArena Vision—1142
GeoBench—52%
SpatialViz-Bench—31.8%

Multilingual DeepSeek-V3.1 leads

DeepSeek-V3.1: 51.6 (#106), Llama 4 Maverick: 42.2 (#195)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1Llama 4 Maverick
LMArena Non-English14001269
LMArena Chinese14691277
LMArena French14471259
LMArena German14111291
LMArena Japanese13781207
LMArena Korean13371203
LMArena Russian14051286
LMArena Spanish14311293

Instruction Following DeepSeek-V3.1 leads

DeepSeek-V3.1: 73.9 (#110), Llama 4 Maverick: 71.7 (#146)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1Llama 4 Maverick
LMArena Instruction Following14001267
IFEval—90.8%

Long Context DeepSeek-V3.1 leads

DeepSeek-V3.1: 36.3 (#232), Llama 4 Maverick: 31.4 (#279)

Long Context benchmarks
BenchmarkDeepSeek-V3.1Llama 4 Maverick
Fiction.LiveBench52.8%46.2%
LMArena Longer Query14221280

Writing & Preference DeepSeek-V3.1 leads

DeepSeek-V3.1: 60.3 (#98), Llama 4 Maverick: 38.8 (#252)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1Llama 4 Maverick
LMArena Text14201287
LMArena Creative Writing14011267
EQ-Bench Creative Writing1436860
LMArena Multi-Turn14081289
Short-Story Creative Writing—62%
WildBench—80%

Frequently asked questions

Is DeepSeek-V3.1 better than Llama 4 Maverick?

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 30.9 on the Noometry Index.

Which is cheaper, DeepSeek-V3.1 or Llama 4 Maverick?

Llama 4 Maverick is cheaper. It lists at $0.19 per million input tokens and $0.65 per million output tokens; DeepSeek-V3.1 lists at $0.25 and $0.95.

Is DeepSeek-V3.1 or Llama 4 Maverick better for coding?

DeepSeek-V3.1 scores higher on coding benchmarks: 40.3 versus 26.6 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3.1 does, with 164K tokens against 128K.

How many benchmarks do DeepSeek-V3.1 and Llama 4 Maverick share?

27 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and Llama 4 Maverick has 54.

Related comparisons

Go deeper