Model comparison

Gemma 3n E4b IT vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 37.3 on the Noometry Index.

Last verified . 14 shared benchmarks.

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 3n E4b IT scores higher in 2 categories and Qwen2.5 Plus 1127 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen2.5 Plus 1127 leads 25.9 to 19.9.
  • Gemma 3n E4b IT has downloadable open weights; the other is API-only.

Side by side

Gemma 3n E4b IT and Qwen2.5 Plus 1127 specifications
Gemma 3n E4b ITQwen2.5 Plus 1127
ProviderGoogleAlibaba (Qwen)
Noometry Index37.338.8
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1814

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Gemma 3n E4b IT: 37.0 (#198), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkGemma 3n E4b ITQwen2.5 Plus 1127
LMArena Coding12681314

Reasoning Qwen2.5 Plus 1127 leads

Gemma 3n E4b IT: 19.9 (#247), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkGemma 3n E4b ITQwen2.5 Plus 1127
LMArena Hard Prompts12841299
Kagi LLM Benchmark31.5%—

Math Qwen2.5 Plus 1127 leads

Gemma 3n E4b IT: 35.1 (#188), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkGemma 3n E4b ITQwen2.5 Plus 1127
LMArena Math12511298

Knowledge Qwen2.5 Plus 1127 leads

Gemma 3n E4b IT: 34.2 (#198), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkGemma 3n E4b ITQwen2.5 Plus 1127
LMArena Expert12461289

Multilingual Gemma 3n E4b IT leads

Gemma 3n E4b IT: 43.4 (#183), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkGemma 3n E4b ITQwen2.5 Plus 1127
LMArena Non-English12851265
LMArena Chinese13091314
LMArena German13111231
LMArena Japanese12721207
LMArena Russian12881271
LMArena French1330—
LMArena Korean1259—
LMArena Spanish1305—

Instruction Following Qwen2.5 Plus 1127 leads

Gemma 3n E4b IT: 66.1 (#210), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkGemma 3n E4b ITQwen2.5 Plus 1127
LMArena Instruction Following12551275

Long Context Too close to call

Gemma 3n E4b IT: 38.7 (#191), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkGemma 3n E4b ITQwen2.5 Plus 1127
LMArena Longer Query12761292

Writing & Preference Too close to call

Gemma 3n E4b IT: 50.1 (#186), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkGemma 3n E4b ITQwen2.5 Plus 1127
LMArena Text13061299
LMArena Creative Writing12871262
LMArena Multi-Turn12761299

Frequently asked questions

Is Gemma 3n E4b IT better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 37.3 on the Noometry Index.

Is Gemma 3n E4b IT or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 37.0 in the Noometry coding category.

How many benchmarks do Gemma 3n E4b IT and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. Gemma 3n E4b IT has 18 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper