Model comparison

Gemma 7B vs Mistral 7B

Gemma 7B is the stronger model overall, scoring 30.0 to 23.0 on the Noometry Index.

Last verified . 27 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Gemma 7B scores higher in 4 categories and Mistral 7B in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 7B leads 31.2 to 8.1.

Side by side

Gemma 7B and Mistral 7B specifications
Gemma 7BMistral 7B
ProviderGoogleMistral AI
Noometry Index30.023.0
Released2024-02-212023-09-27
WeightsOpenOpen
Context window—8K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.25
Results tracked2737

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 7B leads

Gemma 7B: 30.5 (#294), Mistral 7B: 26.4 (#326)

Coding benchmarks
BenchmarkGemma 7BMistral 7B
LMArena Coding10481082
HumanEval+28.7%36%
MBPP+43.4%42.1%
BigCodeBench Instruct—19.5%
BigCodeBench Complete—27.3%

Reasoning Gemma 7B leads

Gemma 7B: 19.9 (#249), Mistral 7B: 13.1 (#336)

Reasoning benchmarks
BenchmarkGemma 7BMistral 7B
LMArena Hard Prompts10421067
Adversarial NLI48.7%47.1%
BIG-Bench Hard55.1%56.1%
Epoch Capabilities Index111.99112.21
HellaSwag82.2%81%
PIQA81.2%83%
WinoGrande79%75.3%
Chess Puzzles—0%
DTBench—42.5%

Math Gemma 7B leads

Gemma 7B: 31.2 (#228), Mistral 7B: 8.1 (#325)

Math benchmarks
BenchmarkGemma 7BMistral 7B
LMArena Math10661085
GSM8K46.4%54.4%
OTIS Mock AIME 2024-2025—0.3%
MATH Level 5—3.7%

Knowledge Gemma 7B leads

Gemma 7B: 27.3 (#252), Mistral 7B: 7.4 (#311)

Knowledge benchmarks
BenchmarkGemma 7BMistral 7B
LMArena Expert10011036
ARC (AI2) Challenge78.3%78.6%
BoolQ83.2%87.4%
MMLU66.1%62.5%
OpenBookQA78.6%79.8%
TriviaQA72.3%75.2%
GPQA Diamond—15.2%

Multilingual Too close to call

Gemma 7B: 25.1 (#287), Mistral 7B: 25.8 (#283)

Multilingual benchmarks
BenchmarkGemma 7BMistral 7B
LMArena Non-English9991012
LMArena Chinese10351009
LMArena French10251037
LMArena Russian9931018
LMArena German—987
LMArena Japanese—878
LMArena Spanish—1026

Instruction Following Mistral 7B leads

Gemma 7B: 51.5 (#295), Mistral 7B: 54.2 (#280)

Instruction Following benchmarks
BenchmarkGemma 7BMistral 7B
LMArena Instruction Following10171060

Long Context Mistral 7B leads

Gemma 7B: 31.1 (#282), Mistral 7B: 32.2 (#271)

Long Context benchmarks
BenchmarkGemma 7BMistral 7B
LMArena Longer Query10221060

Writing & Preference Mistral 7B leads

Gemma 7B: 27.1 (#302), Mistral 7B: 30.7 (#286)

Writing & Preference benchmarks
BenchmarkGemma 7BMistral 7B
LMArena Text10561090
LMArena Creative Writing10241068
LMArena Multi-Turn9631062

Frequently asked questions

Is Gemma 7B better than Mistral 7B?

Gemma 7B is the stronger model overall, scoring 30.0 to 23.0 on the Noometry Index.

Is Gemma 7B or Mistral 7B better for coding?

Gemma 7B scores higher on coding benchmarks: 30.5 versus 26.4 in the Noometry coding category.

How many benchmarks do Gemma 7B and Mistral 7B share?

27 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Mistral 7B has 37.

Related comparisons

Go deeper