Model comparison

DeepSeek-R1-Distill-Qwen-32B vs Gemma 3n E4b IT

Gemma 3n E4b IT is the stronger model overall, scoring 37.3 to 35.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

DeepSeek-R1-Distill-Qwen-32B DeepSeek

35.5

Rank #226 Confirmed

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

Summary

  • The widest gap is in instruction following, where Gemma 3n E4b IT leads 66.1 to 61.6.

Side by side

DeepSeek-R1-Distill-Qwen-32B and Gemma 3n E4b IT specifications
DeepSeek-R1-Distill-Qwen-32BGemma 3n E4b IT
ProviderDeepSeekGoogle
Noometry Index35.537.3
Released2025-01-20—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1418

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek-R1-Distill-Qwen-32B: 36.1 (#212), Gemma 3n E4b IT: 37.0 (#198)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemma 3n E4b IT
BigCodeBench Instruct43.9%—
LiveBench Coding33.7%—
LMArena Coding—1268
BigCodeBench Complete54.9%—

Agentic & Tool Use Not comparable

DeepSeek-R1-Distill-Qwen-32B: 28.1 (#94), Gemma 3n E4b IT: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemma 3n E4b IT
BALROG19.5%—

Reasoning Gemma 3n E4b IT leads

DeepSeek-R1-Distill-Qwen-32B: 18.2 (#284), Gemma 3n E4b IT: 19.9 (#247)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemma 3n E4b IT
Kagi LLM Benchmark—31.5%
Chess Puzzles1%—
LiveBench Reasoning52.3%—
LMArena Hard Prompts—1284
LiveBench Data Analysis45.4%—
Epoch Capabilities Index137.44—
LiveBench45.5%—

Math Too close to call

DeepSeek-R1-Distill-Qwen-32B: 34.5 (#194), Gemma 3n E4b IT: 35.1 (#188)

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemma 3n E4b IT
OTIS Mock AIME 2024-202555.6%—
LiveBench Math59.4%—
LMArena Math—1251

Knowledge DeepSeek-R1-Distill-Qwen-32B leads

DeepSeek-R1-Distill-Qwen-32B: 35.7 (#182), Gemma 3n E4b IT: 34.2 (#198)

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemma 3n E4b IT
GPQA Diamond64.1%—
LMArena Expert—1246

Multilingual Not comparable

DeepSeek-R1-Distill-Qwen-32B: —, Gemma 3n E4b IT: 43.4 (#183)

Multilingual benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemma 3n E4b IT
LMArena Non-English—1285
LMArena Chinese—1309
LMArena French—1330
LMArena German—1311
LMArena Japanese—1272
LMArena Korean—1259
LMArena Russian—1288
LMArena Spanish—1305

Instruction Following Gemma 3n E4b IT leads

DeepSeek-R1-Distill-Qwen-32B: 61.6 (#243), Gemma 3n E4b IT: 66.1 (#210)

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemma 3n E4b IT
LiveBench Instruction Following55.7%—
LMArena Instruction Following—1255

Long Context Not comparable

DeepSeek-R1-Distill-Qwen-32B: —, Gemma 3n E4b IT: 38.7 (#191)

Long Context benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemma 3n E4b IT
LMArena Longer Query—1276

Writing & Preference Too close to call

DeepSeek-R1-Distill-Qwen-32B: 49.6 (#188), Gemma 3n E4b IT: 50.1 (#186)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BGemma 3n E4b IT
LMArena Text—1306
LMArena Creative Writing—1287
LMArena Multi-Turn—1276
LiveBench Language26.8%—

Frequently asked questions

Is DeepSeek-R1-Distill-Qwen-32B better than Gemma 3n E4b IT?

Gemma 3n E4b IT is the stronger model overall, scoring 37.3 to 35.5 on the Noometry Index.

Is DeepSeek-R1-Distill-Qwen-32B or Gemma 3n E4b IT better for coding?

They score almost the same on coding (36.1 vs 37.0); test both on your own repository before choosing.

How many benchmarks do DeepSeek-R1-Distill-Qwen-32B and Gemma 3n E4b IT share?

0 benchmarks have published results for both models. DeepSeek-R1-Distill-Qwen-32B has 14 scored results on Noometry and Gemma 3n E4b IT has 18.

Related comparisons

Go deeper