Model comparison

DeepSeek-V2.5 (Sep 2024) vs Gemma 3n E4b IT

DeepSeek-V2.5 (Sep 2024) and Gemma 3n E4b IT score almost the same on the Noometry Index (37.6 vs 37.3), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

DeepSeek-V2.5 (Sep 2024) DeepSeek

37.6

Rank #200 Confirmed

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

Summary

  • They share 17 benchmarks with published results for both. DeepSeek-V2.5 (Sep 2024) scores higher in 5 categories and Gemma 3n E4b IT in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek-V2.5 (Sep 2024) leads 25.6 to 19.9.

Side by side

DeepSeek-V2.5 (Sep 2024) and Gemma 3n E4b IT specifications
DeepSeek-V2.5 (Sep 2024)Gemma 3n E4b IT
ProviderDeepSeekGoogle
Noometry Index37.637.3
Released2024-09-06—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2218

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 3n E4b IT leads

DeepSeek-V2.5 (Sep 2024): 31.7 (#281), Gemma 3n E4b IT: 37.0 (#198)

Coding benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Gemma 3n E4b IT
LMArena Coding13091268
Aider Polyglot17.8%—
BigCodeBench Instruct48.6%—
BigCodeBench Complete53.2%—
HumanEval+83.5%—
MBPP+74.1%—

Reasoning DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 25.6 (#145), Gemma 3n E4b IT: 19.9 (#247)

Reasoning benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Gemma 3n E4b IT
LMArena Hard Prompts12891284
Kagi LLM Benchmark—31.5%

Math Too close to call

DeepSeek-V2.5 (Sep 2024): 35.9 (#177), Gemma 3n E4b IT: 35.1 (#188)

Math benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Gemma 3n E4b IT
LMArena Math12881251

Knowledge Too close to call

DeepSeek-V2.5 (Sep 2024): 34.8 (#193), Gemma 3n E4b IT: 34.2 (#198)

Knowledge benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Gemma 3n E4b IT
LMArena Expert12661246

Multilingual Too close to call

DeepSeek-V2.5 (Sep 2024): 42.5 (#193), Gemma 3n E4b IT: 43.4 (#183)

Multilingual benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Gemma 3n E4b IT
LMArena Non-English12731285
LMArena Chinese13181309
LMArena French12891330
LMArena German12581311
LMArena Japanese12281272
LMArena Korean12091259
LMArena Russian12891288
LMArena Spanish12481305

Instruction Following DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 67.5 (#194), Gemma 3n E4b IT: 66.1 (#210)

Instruction Following benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Gemma 3n E4b IT
LMArena Instruction Following12801255

Long Context Too close to call

DeepSeek-V2.5 (Sep 2024): 39.5 (#174), Gemma 3n E4b IT: 38.7 (#191)

Long Context benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Gemma 3n E4b IT
LMArena Longer Query13011276

Writing & Preference Too close to call

DeepSeek-V2.5 (Sep 2024): 49.8 (#187), Gemma 3n E4b IT: 50.1 (#186)

Writing & Preference benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Gemma 3n E4b IT
LMArena Text12941306
LMArena Creative Writing12851287
LMArena Multi-Turn12971276

Frequently asked questions

Is DeepSeek-V2.5 (Sep 2024) better than Gemma 3n E4b IT?

DeepSeek-V2.5 (Sep 2024) and Gemma 3n E4b IT score almost the same on the Noometry Index (37.6 vs 37.3), so choose on price, context window or the category you care about most.

Is DeepSeek-V2.5 (Sep 2024) or Gemma 3n E4b IT better for coding?

Gemma 3n E4b IT scores higher on coding benchmarks: 37.0 versus 31.7 in the Noometry coding category.

How many benchmarks do DeepSeek-V2.5 (Sep 2024) and Gemma 3n E4b IT share?

17 benchmarks have published results for both models. DeepSeek-V2.5 (Sep 2024) has 22 scored results on Noometry and Gemma 3n E4b IT has 18.

Related comparisons

Go deeper