Model comparison

Gemma 2 9B vs Mistral Small

Mistral Small is the stronger model overall, scoring 33.4 to 25.9 on the Noometry Index.

Last verified . 30 shared benchmarks.

Gemma 2 9B Google

25.9

Rank #341 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Gemma 2 9B scores higher in 0 categories and Mistral Small in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Small leads 31.0 to 9.7.
  • The biggest single-benchmark swing is LiveBench Reasoning: 15.2% for Gemma 2 9B and 44.8% for Mistral Small.

Side by side

Gemma 2 9B and Mistral Small specifications
Gemma 2 9BMistral Small
ProviderGoogleMistral AI
Noometry Index25.933.4
Released2024-06-242024-02-26
WeightsOpenOpen
Context window—262K
Max output—256K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.60
Results tracked3539

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small leads

Gemma 2 9B: 29.4 (#304), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkGemma 2 9BMistral Small
BigCodeBench Instruct34.7%36.1%
LiveBench Coding22.5%36.2%
LMArena Coding11731362
BigCodeBench Complete40.6%46.6%
SciCode—26.5%
ALE-Bench—497.62

Agentic & Tool Use Not comparable

Gemma 2 9B: —, Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 9BMistral Small
Berkeley Function Calling Leaderboard—37.1%

Reasoning Mistral Small leads

Gemma 2 9B: 15.9 (#309), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkGemma 2 9BMistral Small
LiveBench Reasoning15.2%44.8%
LMArena Hard Prompts11711335
LiveBench Data Analysis36.4%53.7%
LiveBench28.7%44%
Kagi LLM Benchmark—37.8%
CritPt—0%
DTBench—70.9%
LMCA—20.6%
Epoch Capabilities Index119.83—
PIQA83.7%—

Math Mistral Small leads

Gemma 2 9B: 9.9 (#318), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkGemma 2 9BMistral Small
OTIS Mock AIME 2024-20250.6%5.8%
LiveBench Math19.8%39.9%
LMArena Math11831341
MATH Level 521%46.8%
GSM8K84.9%—

Knowledge Mistral Small leads

Gemma 2 9B: 9.7 (#305), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkGemma 2 9BMistral Small
GPQA Diamond27.5%47.5%
LMArena Expert11471291
MMLU72.1%68.7%
Vectara Hallucination Rate—5.1%
BoolQ85.7%—

Multimodal Not comparable

Gemma 2 9B: —, Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkGemma 2 9BMistral Small
LMArena Vision—1142

Multilingual Mistral Small leads

Gemma 2 9B: 36.6 (#238), Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkGemma 2 9BMistral Small
LMArena Non-English11881315
LMArena Chinese11851340
LMArena French11901337
LMArena German11861340
LMArena Japanese11441275
LMArena Korean11371259
LMArena Russian12001324
LMArena Spanish12001346

Instruction Following Mistral Small leads

Gemma 2 9B: 57.6 (#269), Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkGemma 2 9BMistral Small
LiveBench Instruction Following52.6%63.7%
LMArena Instruction Following11781310

Long Context Mistral Small leads

Gemma 2 9B: 36.3 (#233), Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkGemma 2 9BMistral Small
LMArena Longer Query11971327

Writing & Preference Mistral Small leads

Gemma 2 9B: 32.1 (#281), Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkGemma 2 9BMistral Small
LMArena Text12071338
LMArena Creative Writing12061305
LMArena Multi-Turn11931344
LiveBench Language25.5%30.5%
EQ-Bench Creative Writing841—

Frequently asked questions

Is Gemma 2 9B better than Mistral Small?

Mistral Small is the stronger model overall, scoring 33.4 to 25.9 on the Noometry Index.

Is Gemma 2 9B or Mistral Small better for coding?

Mistral Small scores higher on coding benchmarks: 34.0 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2 9B and Mistral Small share?

30 benchmarks have published results for both models. Gemma 2 9B has 35 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper