Model comparison

Gemma 2B vs Magistral Medium

Magistral Medium is the stronger model overall, scoring 35.2 to 29.6 on the Noometry Index.

Last verified . 11 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Gemma 2B scores higher in 1 category and Magistral Medium in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Magistral Medium leads 46.3 to 24.0.

Side by side

Gemma 2B and Magistral Medium specifications
Gemma 2BMagistral Medium
ProviderGoogleMistral AI
Noometry Index29.635.2
Released2024-02-212025-03-17
WeightsOpenOpen
Context window—262K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$5
Results tracked2322

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Gemma 2B: 29.4 (#305), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkGemma 2BMagistral Medium
LMArena Coding10101319
SciCode—39.2%
HumanEval+20.7%—
MBPP+34.1%—

Reasoning Gemma 2B leads

Gemma 2B: 18.8 (#275), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkGemma 2BMagistral Medium
LMArena Hard Prompts9891267
ARC-AGI-2—0%
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
CritPt—0.3%
BIG-Bench Hard35.2%—
Epoch Capabilities Index94.2—
HellaSwag71.4%—
PIQA77.3%—
WinoGrande65.4%—

Math Magistral Medium leads

Gemma 2B: 30.0 (#239), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkGemma 2BMagistral Medium
LMArena Math10091250
GSM8K17.7%—

Knowledge Not comparable

Gemma 2B: —, Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkGemma 2BMagistral Medium
LMArena Expert—1223
ARC (AI2) Challenge42.1%—
BoolQ69.4%—
MMLU42.3%—
TriviaQA53.2%—

Multilingual Magistral Medium leads

Gemma 2B: 23.0 (#294), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkGemma 2BMagistral Medium
LMArena Non-English9581232
LMArena Chinese9861227
LMArena Russian9371224
LMArena French—1267
LMArena German—1248
LMArena Japanese—1175
LMArena Korean—1125
LMArena Spanish—1271

Instruction Following Magistral Medium leads

Gemma 2B: 48.5 (#302), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkGemma 2BMagistral Medium
LMArena Instruction Following9701254

Long Context Magistral Medium leads

Gemma 2B: 29.9 (#291), Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkGemma 2BMagistral Medium
LMArena Longer Query9811295

Writing & Preference Magistral Medium leads

Gemma 2B: 24.0 (#308), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkGemma 2BMagistral Medium
LMArena Text10021255
LMArena Creative Writing9871245
LMArena Multi-Turn9451275

Frequently asked questions

Is Gemma 2B better than Magistral Medium?

Magistral Medium is the stronger model overall, scoring 35.2 to 29.6 on the Noometry Index.

Is Gemma 2B or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2B and Magistral Medium share?

11 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper