Model comparison

Codestral vs Gemma 3 27B

Codestral and Gemma 3 27B score almost the same on the Noometry Index (30.6 vs 30.8), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Gemma 3 27B Google

30.8

Rank #284 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Codestral scores higher in 2 categories and Gemma 3 27B in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Codestral leads 27.3 to 22.5.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 32.5% for Codestral and 40.4% for Gemma 3 27B.
  • Gemma 3 27B is cheaper at $0.08 / $0.16 per million input/output tokens, against $0.30 / $0.90 for Codestral.
  • Codestral accepts more context: 256K tokens versus 131K.
  • Gemma 3 27B has downloadable open weights; the other is API-only.

Side by side

Codestral and Gemma 3 27B specifications
CodestralGemma 3 27B
ProviderMistral AIGoogle
Noometry Index30.630.8
Released2024-05-292025-03-11
WeightsProprietaryOpen
Context window256K131K
Max output8K8K
Input $ / M tokens$0.30$0.08
Output $ / M tokens$0.90$0.16
Results tracked743

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codestral leads

Codestral: 27.3 (#321), Gemma 3 27B: 22.5 (#334)

Coding benchmarks
BenchmarkCodestralGemma 3 27B
Aider Polyglot11.1%4.9%
SciCode—21.2%
BigCodeBench Instruct41.8%—
LiveBench Coding—39.9%
LMArena Coding—1322
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Agentic & Tool Use Not comparable

Codestral: —, Gemma 3 27B: 25.1 (#110)

Agentic & Tool Use benchmarks
BenchmarkCodestralGemma 3 27B
Berkeley Function Calling Leaderboard—29.5%

Reasoning Codestral leads

Codestral: 19.8 (#251), Gemma 3 27B: 16.7 (#301)

Reasoning benchmarks
BenchmarkCodestralGemma 3 27B
Kagi LLM Benchmark32.5%40.4%
CritPt—0%
Chess Puzzles—0%
LiveBench Reasoning—43.8%
LMArena Hard Prompts—1340
DTBench—52.5%
LiveBench Data Analysis—51.5%
LMCA—12.3%
Epoch Capabilities Index—130.04
LiveBench—50%

Math Not comparable

Codestral: —, Gemma 3 27B: 25.9 (#265)

Math benchmarks
BenchmarkCodestralGemma 3 27B
OTIS Mock AIME 2024-2025—22.5%
LiveBench Math—55.4%
LMArena Math—1312
MATH Level 5—74%

Knowledge Not comparable

Codestral: —, Gemma 3 27B: 25.5 (#261)

Knowledge benchmarks
BenchmarkCodestralGemma 3 27B
GPQA Diamond—47.7%
Confabulations—40.3%
Vectara Hallucination Rate—7.4%
LMArena Expert—1304

Multimodal Not comparable

Codestral: —, Gemma 3 27B: 32.6 (#100)

Multimodal benchmarks
BenchmarkCodestralGemma 3 27B
LMArena Vision—1164
GeoBench—52%

Multilingual Not comparable

Codestral: —, Gemma 3 27B: 46.9 (#155)

Multilingual benchmarks
BenchmarkCodestralGemma 3 27B
LMArena Non-English—1334
LMArena Chinese—1346
LMArena French—1368
LMArena German—1362
LMArena Japanese—1287
LMArena Korean—1308
LMArena Russian—1349
LMArena Spanish—1349

Instruction Following Not comparable

Codestral: —, Gemma 3 27B: 70.6 (#160)

Instruction Following benchmarks
BenchmarkCodestralGemma 3 27B
LiveBench Instruction Following—74.9%
LMArena Instruction Following—1321

Long Context Not comparable

Codestral: —, Gemma 3 27B: 27.6 (#293)

Long Context benchmarks
BenchmarkCodestralGemma 3 27B
Fiction.LiveBench—33.3%
LMArena Longer Query—1333

Writing & Preference Not comparable

Codestral: —, Gemma 3 27B: 52.5 (#168)

Writing & Preference benchmarks
BenchmarkCodestralGemma 3 27B
LMArena Text—1358
LMArena Creative Writing—1346
Short-Story Creative Writing—79.9%
EQ-Bench Creative Writing—1266
LMArena Multi-Turn—1345
LiveBench Language—34.6%

Frequently asked questions

Is Codestral better than Gemma 3 27B?

Codestral and Gemma 3 27B score almost the same on the Noometry Index (30.6 vs 30.8), so choose on price, context window or the category you care about most.

Which is cheaper, Codestral or Gemma 3 27B?

Gemma 3 27B is cheaper. It lists at $0.08 per million input tokens and $0.16 per million output tokens; Codestral lists at $0.30 and $0.90.

Is Codestral or Gemma 3 27B better for coding?

Codestral scores higher on coding benchmarks: 27.3 versus 22.5 in the Noometry coding category.

Which has the bigger context window?

Codestral does, with 256K tokens against 131K.

How many benchmarks do Codestral and Gemma 3 27B share?

2 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Gemma 3 27B has 43.

Related comparisons

Go deeper