Model comparison

Gemma 3n E4b IT vs Mistral Large

Gemma 3n E4b IT is the stronger model overall, scoring 37.3 to 31.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 3n E4b IT scores higher in 7 categories and Mistral Large in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 3n E4b IT leads 35.1 to 18.2.

Side by side

Gemma 3n E4b IT and Mistral Large specifications
Gemma 3n E4b ITMistral Large
ProviderGoogleMistral AI
Noometry Index37.331.9
Released—2024-02-26
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1851

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 3n E4b IT leads

Gemma 3n E4b IT: 37.0 (#198), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGemma 3n E4b ITMistral Large
LMArena Coding12681277
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Gemma 3n E4b IT: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGemma 3n E4b ITMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Gemma 3n E4b IT leads

Gemma 3n E4b IT: 19.9 (#247), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGemma 3n E4b ITMistral Large
LMArena Hard Prompts12841257
SimpleBench—22.5%
Kagi LLM Benchmark31.5%—
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Gemma 3n E4b IT leads

Gemma 3n E4b IT: 35.1 (#188), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGemma 3n E4b ITMistral Large
LMArena Math12511262
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Gemma 3n E4b IT leads

Gemma 3n E4b IT: 34.2 (#198), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGemma 3n E4b ITMistral Large
LMArena Expert12461232
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
MMLU—80%

Multilingual Gemma 3n E4b IT leads

Gemma 3n E4b IT: 43.4 (#183), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGemma 3n E4b ITMistral Large
LMArena Non-English12851237
LMArena Chinese13091240
LMArena French13301325
LMArena German13111254
LMArena Japanese12721188
LMArena Korean12591202
LMArena Russian12881257
LMArena Spanish13051268

Instruction Following Mistral Large leads

Gemma 3n E4b IT: 66.1 (#210), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGemma 3n E4b ITMistral Large
LMArena Instruction Following12551249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Too close to call

Gemma 3n E4b IT: 38.7 (#191), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGemma 3n E4b ITMistral Large
LMArena Longer Query12761261

Writing & Preference Gemma 3n E4b IT leads

Gemma 3n E4b IT: 50.1 (#186), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGemma 3n E4b ITMistral Large
LMArena Text13061266
LMArena Creative Writing12871243
LMArena Multi-Turn12761260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Gemma 3n E4b IT better than Mistral Large?

Gemma 3n E4b IT is the stronger model overall, scoring 37.3 to 31.9 on the Noometry Index.

Is Gemma 3n E4b IT or Mistral Large better for coding?

Gemma 3n E4b IT scores higher on coding benchmarks: 37.0 versus 34.3 in the Noometry coding category.

How many benchmarks do Gemma 3n E4b IT and Mistral Large share?

17 benchmarks have published results for both models. Gemma 3n E4b IT has 18 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper