Model comparison

Gemma 1.1 7b IT vs Mistral Medium

Mistral Medium is the stronger model overall, scoring 36.3 to 31.3 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemma 1.1 7b IT Google

31.3

Rank #277 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 1.1 7b IT scores higher in 2 categories and Mistral Medium in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 30.4.

Side by side

Gemma 1.1 7b IT and Mistral Medium specifications
Gemma 1.1 7b ITMistral Medium
ProviderGoogleMistral AI
Noometry Index31.336.3
Released—2023-12-11
WeightsOpenOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked1936

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium leads

Gemma 1.1 7b IT: 31.5 (#284), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkGemma 1.1 7b ITMistral Medium
LMArena Coding10841434
FrontierCode—8%
SciCode—40.2%
WeirdML—43.7%
ALE-Bench—763.98
HumanEval+35.4%—
MBPP+45%—

Agentic & Tool Use Not comparable

Gemma 1.1 7b IT: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkGemma 1.1 7b ITMistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning Mistral Medium leads

Gemma 1.1 7b IT: 20.5 (#238), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkGemma 1.1 7b ITMistral Medium
LMArena Hard Prompts10711426
Kagi LLM Benchmark—50%
CritPt—0%
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Gemma 1.1 7b IT leads

Gemma 1.1 7b IT: 32.0 (#220), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkGemma 1.1 7b ITMistral Medium
LMArena Math11071408
OTIS Mock AIME 2024-2025—32.2%
ProofBench—9%
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Gemma 1.1 7b IT leads

Gemma 1.1 7b IT: 28.3 (#247), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkGemma 1.1 7b ITMistral Medium
LMArena Expert10391408
GPQA Diamond—59.5%
Humanity's Last Exam—4.5%
Vectara Hallucination Rate—22.7%

Multimodal Not comparable

Gemma 1.1 7b IT: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkGemma 1.1 7b ITMistral Medium
LMArena Vision—1172

Multilingual Mistral Medium leads

Gemma 1.1 7b IT: 28.1 (#273), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkGemma 1.1 7b ITMistral Medium
LMArena Non-English10521408
LMArena Chinese10611447
LMArena French10651459
LMArena German10541432
LMArena Japanese9711378
LMArena Korean9881380
LMArena Russian10461411
LMArena Spanish10491433

Instruction Following Mistral Medium leads

Gemma 1.1 7b IT: 54.0 (#283), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkGemma 1.1 7b ITMistral Medium
LMArena Instruction Following10571398

Long Context Mistral Medium leads

Gemma 1.1 7b IT: 32.1 (#272), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkGemma 1.1 7b ITMistral Medium
LMArena Longer Query10561406

Writing & Preference Mistral Medium leads

Gemma 1.1 7b IT: 30.4 (#288), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkGemma 1.1 7b ITMistral Medium
LMArena Text10941424
LMArena Creative Writing10601391
LMArena Multi-Turn10401418
Short-Story Creative Writing—77.3%

Frequently asked questions

Is Gemma 1.1 7b IT better than Mistral Medium?

Mistral Medium is the stronger model overall, scoring 36.3 to 31.3 on the Noometry Index.

Is Gemma 1.1 7b IT or Mistral Medium better for coding?

Mistral Medium scores higher on coding benchmarks: 34.2 versus 31.5 in the Noometry coding category.

How many benchmarks do Gemma 1.1 7b IT and Mistral Medium share?

17 benchmarks have published results for both models. Gemma 1.1 7b IT has 19 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper