Model comparison

Codestral vs Granite 4.2 3b

Granite 4.2 3b is the stronger model overall, scoring 39.4 to 30.6 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Granite 4.2 3b IBM

39.4

Rank #169 Confirmed

Summary

  • The widest gap is in coding, where Granite 4.2 3b leads 40.0 to 27.3.
  • Granite 4.2 3b has downloadable open weights; the other is API-only.

Side by side

Codestral and Granite 4.2 3b specifications
CodestralGranite 4.2 3b
ProviderMistral AIIBM
Noometry Index30.639.4
Released2024-05-29—
WeightsProprietaryOpen
Context window256K—
Max output8K—
Input $ / M tokens$0.30—
Output $ / M tokens$0.90—
Results tracked711

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.2 3b leads

Codestral: 27.3 (#321), Granite 4.2 3b: 40.0 (#151)

Coding benchmarks
BenchmarkCodestralGranite 4.2 3b
Aider Polyglot11.1%—
BigCodeBench Instruct41.8%—
LMArena Coding—1361
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Reasoning Granite 4.2 3b leads

Codestral: 19.8 (#251), Granite 4.2 3b: 26.0 (#138)

Reasoning benchmarks
BenchmarkCodestralGranite 4.2 3b
Kagi LLM Benchmark32.5%—
LMArena Hard Prompts—1306

Knowledge Not comparable

Codestral: —, Granite 4.2 3b: 36.3 (#171)

Knowledge benchmarks
BenchmarkCodestralGranite 4.2 3b
LMArena Expert—1315

Multilingual Not comparable

Codestral: —, Granite 4.2 3b: 42.1 (#198)

Multilingual benchmarks
BenchmarkCodestralGranite 4.2 3b
LMArena Non-English—1268
LMArena Chinese—1269
LMArena Russian—1249

Instruction Following Not comparable

Codestral: —, Granite 4.2 3b: 67.1 (#200)

Instruction Following benchmarks
BenchmarkCodestralGranite 4.2 3b
LMArena Instruction Following—1273

Long Context Not comparable

Codestral: —, Granite 4.2 3b: 39.2 (#185)

Long Context benchmarks
BenchmarkCodestralGranite 4.2 3b
LMArena Longer Query—1291

Writing & Preference Not comparable

Codestral: —, Granite 4.2 3b: 47.2 (#212)

Writing & Preference benchmarks
BenchmarkCodestralGranite 4.2 3b
LMArena Text—1293
LMArena Creative Writing—1205
LMArena Multi-Turn—1290

Frequently asked questions

Is Codestral better than Granite 4.2 3b?

Granite 4.2 3b is the stronger model overall, scoring 39.4 to 30.6 on the Noometry Index.

Is Codestral or Granite 4.2 3b better for coding?

Granite 4.2 3b scores higher on coding benchmarks: 40.0 versus 27.3 in the Noometry coding category.

How many benchmarks do Codestral and Granite 4.2 3b share?

0 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Granite 4.2 3b has 11.

Related comparisons

Go deeper