Model comparison

Gemma 2 2b IT vs Mistral Medium

Mistral Medium is the stronger model overall, scoring 36.3 to 33.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemma 2 2b IT Google

33.1

Rank #248 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 2 2b IT scores higher in 2 categories and Mistral Medium in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 36.5.

Side by side

Gemma 2 2b IT and Mistral Medium specifications
Gemma 2 2b ITMistral Medium
ProviderGoogleMistral AI
Noometry Index33.136.3
Released—2023-12-11
WeightsOpenOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked1736

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium leads

Gemma 2 2b IT: 32.3 (#273), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkGemma 2 2b ITMistral Medium
LMArena Coding11121434
FrontierCode—8%
SciCode—40.2%
WeirdML—43.7%
ALE-Bench—763.98

Agentic & Tool Use Not comparable

Gemma 2 2b IT: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 2b ITMistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning Mistral Medium leads

Gemma 2 2b IT: 21.4 (#224), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkGemma 2 2b ITMistral Medium
LMArena Hard Prompts11131426
Kagi LLM Benchmark—50%
CritPt—0%
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Gemma 2 2b IT leads

Gemma 2 2b IT: 32.6 (#212), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkGemma 2 2b ITMistral Medium
LMArena Math11351408
OTIS Mock AIME 2024-2025—32.2%
ProofBench—9%
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Gemma 2 2b IT leads

Gemma 2 2b IT: 29.9 (#231), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkGemma 2 2b ITMistral Medium
LMArena Expert10961408
GPQA Diamond—59.5%
Humanity's Last Exam—4.5%
Vectara Hallucination Rate—22.7%

Multimodal Not comparable

Gemma 2 2b IT: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkGemma 2 2b ITMistral Medium
LMArena Vision—1172

Multilingual Mistral Medium leads

Gemma 2 2b IT: 32.3 (#257), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkGemma 2 2b ITMistral Medium
LMArena Non-English11211408
LMArena Chinese11321447
LMArena French11571459
LMArena German11141432
LMArena Japanese10831378
LMArena Korean10551380
LMArena Russian11181411
LMArena Spanish11401433

Instruction Following Mistral Medium leads

Gemma 2 2b IT: 57.8 (#263), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkGemma 2 2b ITMistral Medium
LMArena Instruction Following11181398

Long Context Mistral Medium leads

Gemma 2 2b IT: 34.3 (#250), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkGemma 2 2b ITMistral Medium
LMArena Longer Query11301406

Writing & Preference Mistral Medium leads

Gemma 2 2b IT: 36.5 (#263), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkGemma 2 2b ITMistral Medium
LMArena Text11561424
LMArena Creative Writing11471391
LMArena Multi-Turn11181418
Short-Story Creative Writing—77.3%

Frequently asked questions

Is Gemma 2 2b IT better than Mistral Medium?

Mistral Medium is the stronger model overall, scoring 36.3 to 33.1 on the Noometry Index.

Is Gemma 2 2b IT or Mistral Medium better for coding?

Mistral Medium scores higher on coding benchmarks: 34.2 versus 32.3 in the Noometry coding category.

How many benchmarks do Gemma 2 2b IT and Mistral Medium share?

17 benchmarks have published results for both models. Gemma 2 2b IT has 17 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper