Model comparison

Gemma 3n E4b IT vs Qwen3 32B

Qwen3 32B is the stronger model overall, scoring 39.2 to 37.3 on the Noometry Index.

Last verified . 14 shared benchmarks.

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 3n E4b IT scores higher in 0 categories and Qwen3 32B in 8 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 32B leads 40.0 to 34.2.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 31.5% for Gemma 3n E4b IT and 54.9% for Qwen3 32B.

Side by side

Gemma 3n E4b IT and Qwen3 32B specifications
Gemma 3n E4b ITQwen3 32B
ProviderGoogleAlibaba (Qwen)
Noometry Index37.339.2
Released—2025-04
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$0.70
Output $ / M tokens—$2.80
Results tracked1826

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 3n E4b IT: 37.0 (#198), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkGemma 3n E4b ITQwen3 32B
LMArena Coding12681358
Aider Polyglot—40%
SciCode—35.4%

Agentic & Tool Use Not comparable

Gemma 3n E4b IT: —, Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkGemma 3n E4b ITQwen3 32B
Berkeley Function Calling Leaderboard—48.7%

Reasoning Too close to call

Gemma 3n E4b IT: 19.9 (#247), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkGemma 3n E4b ITQwen3 32B
Kagi LLM Benchmark31.5%54.9%
LMArena Hard Prompts12841334
CritPt—0.3%
Chess Puzzles—5%
DTBench—67.5%
LMCA—17.3%
Epoch Capabilities Index—138.51

Math Qwen3 32B leads

Gemma 3n E4b IT: 35.1 (#188), Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkGemma 3n E4b ITQwen3 32B
LMArena Math12511399
OTIS Mock AIME 2024-2025—66.9%

Knowledge Qwen3 32B leads

Gemma 3n E4b IT: 34.2 (#198), Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkGemma 3n E4b ITQwen3 32B
LMArena Expert12461362
GPQA Diamond—65.7%
Vectara Hallucination Rate—5.9%

Multilingual Qwen3 32B leads

Gemma 3n E4b IT: 43.4 (#183), Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkGemma 3n E4b ITQwen3 32B
LMArena Non-English12851317
LMArena Chinese13091357
LMArena German13111341
LMArena Russian12881311
LMArena French1330—
LMArena Japanese1272—
LMArena Korean1259—
LMArena Spanish1305—

Instruction Following Qwen3 32B leads

Gemma 3n E4b IT: 66.1 (#210), Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkGemma 3n E4b ITQwen3 32B
LMArena Instruction Following12551305

Long Context Qwen3 32B leads

Gemma 3n E4b IT: 38.7 (#191), Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkGemma 3n E4b ITQwen3 32B
LMArena Longer Query12761327
Fiction.LiveBench—74.2%

Writing & Preference Qwen3 32B leads

Gemma 3n E4b IT: 50.1 (#186), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkGemma 3n E4b ITQwen3 32B
LMArena Text13061340
LMArena Creative Writing12871297
LMArena Multi-Turn12761331

Frequently asked questions

Is Gemma 3n E4b IT better than Qwen3 32B?

Qwen3 32B is the stronger model overall, scoring 39.2 to 37.3 on the Noometry Index.

Is Gemma 3n E4b IT or Qwen3 32B better for coding?

They score almost the same on coding (37.0 vs 37.7); test both on your own repository before choosing.

How many benchmarks do Gemma 3n E4b IT and Qwen3 32B share?

14 benchmarks have published results for both models. Gemma 3n E4b IT has 18 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper