Model comparison

Gemma 3n E4b IT vs Magistral Small

Gemma 3n E4b IT is the stronger model overall, scoring 37.3 to 30.2 on the Noometry Index.

Last verified . 1 shared benchmarks.

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Summary

  • They share 1 benchmark with published results for both. Gemma 3n E4b IT scores higher in 3 categories and Magistral Small in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemma 3n E4b IT leads 19.9 to 6.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 31.5% for Gemma 3n E4b IT and 6.3% for Magistral Small.

Side by side

Gemma 3n E4b IT and Magistral Small specifications
Gemma 3n E4b ITMagistral Small
ProviderGoogleMistral AI
Noometry Index37.330.2
Released—2025-06-10
WeightsOpenOpen
Context window—128K
Max output—40K
Input $ / M tokens—$0.50
Output $ / M tokens—$1.50
Results tracked1810

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Small leads

Gemma 3n E4b IT: 37.0 (#198), Magistral Small: 38.4 (#176)

Coding benchmarks
BenchmarkGemma 3n E4b ITMagistral Small
SciCode—35.2%
LMArena Coding1268—

Reasoning Gemma 3n E4b IT leads

Gemma 3n E4b IT: 19.9 (#247), Magistral Small: 6.8 (#350)

Reasoning benchmarks
BenchmarkGemma 3n E4b ITMagistral Small
Kagi LLM Benchmark31.5%6.3%
ARC-AGI-2—0%
ARC-AGI-1—5%
CritPt—0.3%
Chess Puzzles—3%
LMArena Hard Prompts1284—
DTBench—61.3%
Epoch Capabilities Index—133.19

Math Gemma 3n E4b IT leads

Gemma 3n E4b IT: 35.1 (#188), Magistral Small: 26.2 (#261)

Math benchmarks
BenchmarkGemma 3n E4b ITMagistral Small
OTIS Mock AIME 2024-2025—30%
LMArena Math1251—

Knowledge Gemma 3n E4b IT leads

Gemma 3n E4b IT: 34.2 (#198), Magistral Small: 30.9 (#223)

Knowledge benchmarks
BenchmarkGemma 3n E4b ITMagistral Small
GPQA Diamond—56.1%
LMArena Expert1246—

Multilingual Not comparable

Gemma 3n E4b IT: 43.4 (#183), Magistral Small: —

Multilingual benchmarks
BenchmarkGemma 3n E4b ITMagistral Small
LMArena Non-English1285—
LMArena Chinese1309—
LMArena French1330—
LMArena German1311—
LMArena Japanese1272—
LMArena Korean1259—
LMArena Russian1288—
LMArena Spanish1305—

Instruction Following Not comparable

Gemma 3n E4b IT: 66.1 (#210), Magistral Small: —

Instruction Following benchmarks
BenchmarkGemma 3n E4b ITMagistral Small
LMArena Instruction Following1255—

Long Context Not comparable

Gemma 3n E4b IT: 38.7 (#191), Magistral Small: —

Long Context benchmarks
BenchmarkGemma 3n E4b ITMagistral Small
LMArena Longer Query1276—

Writing & Preference Not comparable

Gemma 3n E4b IT: 50.1 (#186), Magistral Small: —

Writing & Preference benchmarks
BenchmarkGemma 3n E4b ITMagistral Small
LMArena Text1306—
LMArena Creative Writing1287—
LMArena Multi-Turn1276—

Frequently asked questions

Is Gemma 3n E4b IT better than Magistral Small?

Gemma 3n E4b IT is the stronger model overall, scoring 37.3 to 30.2 on the Noometry Index.

Is Gemma 3n E4b IT or Magistral Small better for coding?

Magistral Small scores higher on coding benchmarks: 38.4 versus 37.0 in the Noometry coding category.

How many benchmarks do Gemma 3n E4b IT and Magistral Small share?

1 benchmark has published results for both models. Gemma 3n E4b IT has 18 scored results on Noometry and Magistral Small has 10.

Related comparisons

Go deeper