Model comparison

Gemma 1.1 2b IT vs Mistral Medium

Mistral Medium is the stronger model overall, scoring 36.3 to 29.3 on the Noometry Index.

Last verified . 14 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 2 categories and Mistral Medium in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 25.1.

Side by side

Gemma 1.1 2b IT and Mistral Medium specifications
Gemma 1.1 2b ITMistral Medium
ProviderGoogleMistral AI
Noometry Index29.336.3
Released—2023-12-11
WeightsOpenOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked1636

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium leads

Gemma 1.1 2b IT: 30.1 (#299), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium
LMArena Coding10341434
FrontierCode—8%
SciCode—40.2%
WeirdML—43.7%
ALE-Bench—763.98
HumanEval+17.7%—
MBPP+23.3%—

Agentic & Tool Use Not comparable

Gemma 1.1 2b IT: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning Mistral Medium leads

Gemma 1.1 2b IT: 19.1 (#270), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium
LMArena Hard Prompts10051426
Kagi LLM Benchmark—50%
CritPt—0%
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 30.8 (#232), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium
LMArena Math10471408
OTIS Mock AIME 2024-2025—32.2%
ProofBench—9%
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 26.5 (#258), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium
LMArena Expert9701408
GPQA Diamond—59.5%
Humanity's Last Exam—4.5%
Vectara Hallucination Rate—22.7%

Multimodal Not comparable

Gemma 1.1 2b IT: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium
LMArena Vision—1172

Multilingual Mistral Medium leads

Gemma 1.1 2b IT: 24.6 (#289), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium
LMArena Non-English9881408
LMArena Chinese10121447
LMArena German9441432
LMArena Korean8991380
LMArena Russian9901411
LMArena French—1459
LMArena Japanese—1378
LMArena Spanish—1433

Instruction Following Mistral Medium leads

Gemma 1.1 2b IT: 49.9 (#299), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium
LMArena Instruction Following9921398

Long Context Mistral Medium leads

Gemma 1.1 2b IT: 30.6 (#286), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium
LMArena Longer Query10031406

Writing & Preference Mistral Medium leads

Gemma 1.1 2b IT: 25.1 (#306), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium
LMArena Text10221424
LMArena Creative Writing9981391
LMArena Multi-Turn9591418
Short-Story Creative Writing—77.3%

Frequently asked questions

Is Gemma 1.1 2b IT better than Mistral Medium?

Mistral Medium is the stronger model overall, scoring 36.3 to 29.3 on the Noometry Index.

Is Gemma 1.1 2b IT or Mistral Medium better for coding?

Mistral Medium scores higher on coding benchmarks: 34.2 versus 30.1 in the Noometry coding category.

How many benchmarks do Gemma 1.1 2b IT and Mistral Medium share?

14 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper