Model comparison

Gemma 1.1 2b IT vs Mistral Medium 3.5

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 29.3 on the Noometry Index.

Last verified . 14 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 1 category and Mistral Medium 3.5 in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium 3.5 leads 58.5 to 25.1.

Side by side

Gemma 1.1 2b IT and Mistral Medium 3.5 specifications
Gemma 1.1 2b ITMistral Medium 3.5
ProviderGoogleMistral AI
Noometry Index29.340.2
Released——
WeightsOpenOpen
Context window—262K
Max output—210K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked1622

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium 3.5 leads

Gemma 1.1 2b IT: 30.1 (#299), Mistral Medium 3.5: 36.0 (#213)

Coding benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium 3.5
LMArena Coding10341461
LMArena WebDev—1264
HumanEval+17.7%—
MBPP+23.3%—

Reasoning Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 19.1 (#270), Mistral Medium 3.5: 17.3 (#295)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium 3.5
LMArena Hard Prompts10051436
Kagi LLM Benchmark—41.4%
NYT Connections (extended)—12.9%
Epoch Capabilities Index—141.35

Math Mistral Medium 3.5 leads

Gemma 1.1 2b IT: 30.8 (#232), Mistral Medium 3.5: 39.1 (#113)

Math benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium 3.5
LMArena Math10471431

Knowledge Mistral Medium 3.5 leads

Gemma 1.1 2b IT: 26.5 (#258), Mistral Medium 3.5: 40.0 (#126)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium 3.5
LMArena Expert9701432

Multimodal Not comparable

Gemma 1.1 2b IT: —, Mistral Medium 3.5: 38.3 (#65)

Multimodal benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium 3.5
LMArena Vision—1223

Multilingual Mistral Medium 3.5 leads

Gemma 1.1 2b IT: 24.6 (#289), Mistral Medium 3.5: 51.9 (#100)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium 3.5
LMArena Non-English9881404
LMArena Chinese10121442
LMArena German9441451
LMArena Korean8991385
LMArena Russian9901395
LMArena French—1448
LMArena Spanish—1409

Instruction Following Mistral Medium 3.5 leads

Gemma 1.1 2b IT: 49.9 (#299), Mistral Medium 3.5: 74.6 (#90)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium 3.5
LMArena Instruction Following9921415

Long Context Mistral Medium 3.5 leads

Gemma 1.1 2b IT: 30.6 (#286), Mistral Medium 3.5: 43.2 (#103)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium 3.5
LMArena Longer Query10031415

Writing & Preference Mistral Medium 3.5 leads

Gemma 1.1 2b IT: 25.1 (#306), Mistral Medium 3.5: 58.5 (#117)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITMistral Medium 3.5
LMArena Text10221421
LMArena Creative Writing9981374
LMArena Multi-Turn9591423
EQ-Bench 4—993

Frequently asked questions

Is Gemma 1.1 2b IT better than Mistral Medium 3.5?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 29.3 on the Noometry Index.

Is Gemma 1.1 2b IT or Mistral Medium 3.5 better for coding?

Mistral Medium 3.5 scores higher on coding benchmarks: 36.0 versus 30.1 in the Noometry coding category.

How many benchmarks do Gemma 1.1 2b IT and Mistral Medium 3.5 share?

14 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Mistral Medium 3.5 has 22.

Related comparisons

Go deeper