Model comparison

Gemma 3n E4b IT vs Olmo 3.1 32b Think

Gemma 3n E4b IT and Olmo 3.1 32b Think score almost the same on the Noometry Index (37.3 vs 37.9), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Gemma 3n E4b IT scores higher in 4 categories and Olmo 3.1 32b Think in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Olmo 3.1 32b Think leads 25.2 to 19.9.

Side by side

Gemma 3n E4b IT and Olmo 3.1 32b Think specifications
Gemma 3n E4b ITOlmo 3.1 32b Think
ProviderGoogleAllen Institute for AI (Ai2)
Noometry Index37.337.9
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1815

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 3n E4b IT: 37.0 (#198), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkGemma 3n E4b ITOlmo 3.1 32b Think
LMArena Coding12681291

Reasoning Olmo 3.1 32b Think leads

Gemma 3n E4b IT: 19.9 (#247), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkGemma 3n E4b ITOlmo 3.1 32b Think
LMArena Hard Prompts12841272
Kagi LLM Benchmark31.5%—

Math Olmo 3.1 32b Think leads

Gemma 3n E4b IT: 35.1 (#188), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkGemma 3n E4b ITOlmo 3.1 32b Think
LMArena Math12511305

Knowledge Olmo 3.1 32b Think leads

Gemma 3n E4b IT: 34.2 (#198), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkGemma 3n E4b ITOlmo 3.1 32b Think
LMArena Expert12461295

Multilingual Gemma 3n E4b IT leads

Gemma 3n E4b IT: 43.4 (#183), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkGemma 3n E4b ITOlmo 3.1 32b Think
LMArena Non-English12851209
LMArena Chinese13091242
LMArena French13301260
LMArena German13111262
LMArena Russian12881193
LMArena Spanish13051289
LMArena Japanese1272—
LMArena Korean1259—

Instruction Following Too close to call

Gemma 3n E4b IT: 66.1 (#210), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkGemma 3n E4b ITOlmo 3.1 32b Think
LMArena Instruction Following12551247

Long Context Too close to call

Gemma 3n E4b IT: 38.7 (#191), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkGemma 3n E4b ITOlmo 3.1 32b Think
LMArena Longer Query12761272

Writing & Preference Gemma 3n E4b IT leads

Gemma 3n E4b IT: 50.1 (#186), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkGemma 3n E4b ITOlmo 3.1 32b Think
LMArena Text13061272
LMArena Creative Writing12871226
LMArena Multi-Turn12761252

Frequently asked questions

Is Gemma 3n E4b IT better than Olmo 3.1 32b Think?

Gemma 3n E4b IT and Olmo 3.1 32b Think score almost the same on the Noometry Index (37.3 vs 37.9), so choose on price, context window or the category you care about most.

Is Gemma 3n E4b IT or Olmo 3.1 32b Think better for coding?

They score almost the same on coding (37.0 vs 37.7); test both on your own repository before choosing.

How many benchmarks do Gemma 3n E4b IT and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. Gemma 3n E4b IT has 18 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper