Model comparison

Gemma 3 4B vs Mistral Nemo

Gemma 3 4B is the stronger model overall, scoring 28.1 to 26.4 on the Noometry Index.

Last verified . 5 shared benchmarks.

Gemma 3 4B Google

28.1

Rank #326 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Gemma 3 4B scores higher in 1 category and Mistral Nemo in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemma 3 4B leads 42.0 to 28.5.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 19.6% for Gemma 3 4B and 27.6% for Mistral Nemo.
  • Gemma 3 4B is cheaper at $0.04 / $0.08 per million input/output tokens, against $0.15 / $0.15 for Mistral Nemo.
  • Gemma 3 4B accepts more context: 131K tokens versus 128K.

Side by side

Gemma 3 4B and Mistral Nemo specifications
Gemma 3 4BMistral Nemo
ProviderGoogleMistral AI
Noometry Index28.126.4
Released2025-03-122024-07-01
WeightsOpenOpen
Context window131K128K
Max output4K128K
Input $ / M tokens$0.04$0.15
Output $ / M tokens$0.08$0.15
Results tracked2210

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Gemma 3 4B: 35.9 (#215), Mistral Nemo: —

Coding benchmarks
BenchmarkGemma 3 4BMistral Nemo
LMArena Coding1230—

Agentic & Tool Use Mistral Nemo leads

Gemma 3 4B: 20.9 (#142), Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkGemma 3 4BMistral Nemo
Berkeley Function Calling Leaderboard19.6%27.6%
BALROG—17.6%

Reasoning Mistral Nemo leads

Gemma 3 4B: 13.2 (#335), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkGemma 3 4BMistral Nemo
DTBench50.9%48.6%
Epoch Capabilities Index116.02118.68
Kagi LLM Benchmark25.2%—
Chess Puzzles0%—
LMArena Hard Prompts1253—
LMCA2.8%—
PIQA—83.5%

Math Mistral Nemo leads

Gemma 3 4B: 16.8 (#292), Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkGemma 3 4BMistral Nemo
OTIS Mock AIME 2024-20257.5%—
LMArena Math1239—
MATH Level 5—10.8%
GSM8K—84.2%

Knowledge Too close to call

Gemma 3 4B: 11.8 (#299), Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkGemma 3 4BMistral Nemo
GPQA Diamond23.2%29.9%
Vectara Hallucination Rate6.4%—
LMArena Expert1223—
BoolQ—82.5%

Multilingual Not comparable

Gemma 3 4B: 42.5 (#194), Mistral Nemo: —

Multilingual benchmarks
BenchmarkGemma 3 4BMistral Nemo
LMArena Non-English1273—
LMArena German1281—
LMArena Russian1294—

Instruction Following Not comparable

Gemma 3 4B: 65.2 (#225), Mistral Nemo: —

Instruction Following benchmarks
BenchmarkGemma 3 4BMistral Nemo
LMArena Instruction Following1239—

Long Context Not comparable

Gemma 3 4B: 38.7 (#194), Mistral Nemo: —

Long Context benchmarks
BenchmarkGemma 3 4BMistral Nemo
LMArena Longer Query1273—

Writing & Preference Gemma 3 4B leads

Gemma 3 4B: 42.0 (#239), Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkGemma 3 4BMistral Nemo
EQ-Bench Creative Writing1068881
LMArena Text1291—
LMArena Creative Writing1271—
LMArena Multi-Turn1255—

Frequently asked questions

Is Gemma 3 4B better than Mistral Nemo?

Gemma 3 4B is the stronger model overall, scoring 28.1 to 26.4 on the Noometry Index.

Which is cheaper, Gemma 3 4B or Mistral Nemo?

Gemma 3 4B is cheaper. It lists at $0.04 per million input tokens and $0.08 per million output tokens; Mistral Nemo lists at $0.15 and $0.15.

Which has the bigger context window?

Gemma 3 4B does, with 131K tokens against 128K.

How many benchmarks do Gemma 3 4B and Mistral Nemo share?

5 benchmarks have published results for both models. Gemma 3 4B has 22 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper