Model comparison

Gemma 3 27B vs Mistral Small 3.1

Gemma 3 27B and Mistral Small 3.1 score almost the same on the Noometry Index (30.8 vs 31.7), so choose on price, context window or the category you care about most.

Last verified . 23 shared benchmarks.

Gemma 3 27B Google

30.8

Rank #284 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Gemma 3 27B scores higher in 5 categories and Mistral Small 3.1 in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Mistral Small 3.1 leads 38.3 to 22.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 22.5% for Gemma 3 27B and 3.9% for Mistral Small 3.1.
  • Gemma 3 27B is cheaper at $0.08 / $0.16 per million input/output tokens, against $0.35 / $0.56 for Mistral Small 3.1.
  • Gemma 3 27B accepts more context: 131K tokens versus 128K.

Side by side

Gemma 3 27B and Mistral Small 3.1 specifications
Gemma 3 27BMistral Small 3.1
ProviderGoogleMistral AI
Noometry Index30.831.7
Released2025-03-112025-03-17
WeightsOpenOpen
Context window131K128K
Max output8K102K
Input $ / M tokens$0.08$0.35
Output $ / M tokens$0.16$0.56
Results tracked4328

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3.1 leads

Gemma 3 27B: 22.5 (#334), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkGemma 3 27BMistral Small 3.1
LMArena Coding13221309
Aider Polyglot4.9%—
SciCode21.2%—
LiveBench Coding39.9%—

Agentic & Tool Use Not comparable

Gemma 3 27B: 25.1 (#110), Mistral Small 3.1: —

Agentic & Tool Use benchmarks
BenchmarkGemma 3 27BMistral Small 3.1
Berkeley Function Calling Leaderboard29.5%—

Reasoning Mistral Small 3.1 leads

Gemma 3 27B: 16.7 (#301), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkGemma 3 27BMistral Small 3.1
Chess Puzzles0%1%
LMArena Hard Prompts13401278
Epoch Capabilities Index130.04127.48
Kagi LLM Benchmark40.4%—
CritPt0%—
LiveBench Reasoning43.8%—
DTBench52.5%—
LiveBench Data Analysis51.5%—
LMCA12.3%—
LiveBench50%—

Math Gemma 3 27B leads

Gemma 3 27B: 25.9 (#265), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkGemma 3 27BMistral Small 3.1
OTIS Mock AIME 2024-202522.5%3.9%
LMArena Math13121262
Omni-MATH—24.8%
LiveBench Math55.4%—
MATH Level 574%—

Knowledge Gemma 3 27B leads

Gemma 3 27B: 25.5 (#261), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkGemma 3 27BMistral Small 3.1
GPQA Diamond47.7%41.9%
LMArena Expert13041257
MMLU-Pro—61%
Confabulations40.3%—
Vectara Hallucination Rate7.4%—
GPQA (HELM)—39.2%

Multimodal Too close to call

Gemma 3 27B: 32.6 (#100), Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkGemma 3 27BMistral Small 3.1
LMArena Vision11641136
GeoBench52%—

Multilingual Gemma 3 27B leads

Gemma 3 27B: 46.9 (#155), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkGemma 3 27BMistral Small 3.1
LMArena Non-English13341255
LMArena Chinese13461253
LMArena French13681273
LMArena German13621266
LMArena Japanese12871208
LMArena Korean13081206
LMArena Russian13491263
LMArena Spanish13491283

Instruction Following Gemma 3 27B leads

Gemma 3 27B: 70.6 (#160), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkGemma 3 27BMistral Small 3.1
LMArena Instruction Following13211264
LiveBench Instruction Following74.9%—
IFEval—75%

Long Context Mistral Small 3.1 leads

Gemma 3 27B: 27.6 (#293), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkGemma 3 27BMistral Small 3.1
LMArena Longer Query13331299
Fiction.LiveBench33.3%—

Writing & Preference Gemma 3 27B leads

Gemma 3 27B: 52.5 (#168), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkGemma 3 27BMistral Small 3.1
LMArena Text13581277
LMArena Creative Writing13461253
EQ-Bench Creative Writing1266761
LMArena Multi-Turn13451270
Short-Story Creative Writing79.9%—
WildBench—78.8%
LiveBench Language34.6%—

Frequently asked questions

Is Gemma 3 27B better than Mistral Small 3.1?

Gemma 3 27B and Mistral Small 3.1 score almost the same on the Noometry Index (30.8 vs 31.7), so choose on price, context window or the category you care about most.

Which is cheaper, Gemma 3 27B or Mistral Small 3.1?

Gemma 3 27B is cheaper. It lists at $0.08 per million input tokens and $0.16 per million output tokens; Mistral Small 3.1 lists at $0.35 and $0.56.

Is Gemma 3 27B or Mistral Small 3.1 better for coding?

Mistral Small 3.1 scores higher on coding benchmarks: 38.3 versus 22.5 in the Noometry coding category.

Which has the bigger context window?

Gemma 3 27B does, with 131K tokens against 128K.

How many benchmarks do Gemma 3 27B and Mistral Small 3.1 share?

23 benchmarks have published results for both models. Gemma 3 27B has 43 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper