Model comparison

Gemma 7B vs Mistral Large

Mistral Large is the stronger model overall, scoring 31.9 to 30.0 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 7B scores higher in 2 categories and Mistral Large in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Mistral Large leads 67.9 to 51.5.

Side by side

Gemma 7B and Mistral Large specifications
Gemma 7BMistral Large
ProviderGoogleMistral AI
Noometry Index30.031.9
Released2024-02-212024-02-26
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked2751

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Gemma 7B: 30.5 (#294), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGemma 7BMistral Large
LMArena Coding10481277
HumanEval+28.7%62.2%
MBPP+43.4%59.5%
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7

Agentic & Tool Use Not comparable

Gemma 7B: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGemma 7BMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Gemma 7B leads

Gemma 7B: 19.9 (#249), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGemma 7BMistral Large
LMArena Hard Prompts10421257
Epoch Capabilities Index111.99128.52
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Adversarial NLI48.7%—
BIG-Bench Hard55.1%—
ForecastBench—57.1
HellaSwag82.2%—
LiveBench—48.4%
PIQA81.2%—
WinoGrande79%—

Math Gemma 7B leads

Gemma 7B: 31.2 (#228), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGemma 7BMistral Large
LMArena Math10661262
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%
GSM8K46.4%—

Knowledge Mistral Large leads

Gemma 7B: 27.3 (#252), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGemma 7BMistral Large
LMArena Expert10011232
MMLU66.1%80%
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
ARC (AI2) Challenge78.3%—
BoolQ83.2%—
OpenBookQA78.6%—
TriviaQA72.3%—

Multilingual Mistral Large leads

Gemma 7B: 25.1 (#287), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGemma 7BMistral Large
LMArena Non-English9991237
LMArena Chinese10351240
LMArena French10251325
LMArena Russian9931257
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Spanish—1268

Instruction Following Mistral Large leads

Gemma 7B: 51.5 (#295), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGemma 7BMistral Large
LMArena Instruction Following10171249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Mistral Large leads

Gemma 7B: 31.1 (#282), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGemma 7BMistral Large
LMArena Longer Query10221261

Writing & Preference Mistral Large leads

Gemma 7B: 27.1 (#302), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGemma 7BMistral Large
LMArena Text10561266
LMArena Creative Writing10241243
LMArena Multi-Turn9631260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Gemma 7B better than Mistral Large?

Mistral Large is the stronger model overall, scoring 31.9 to 30.0 on the Noometry Index.

Is Gemma 7B or Mistral Large better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 30.5 in the Noometry coding category.

How many benchmarks do Gemma 7B and Mistral Large share?

17 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper