Model comparison

Granite 4.0 Micro vs Mistral Large

Mistral Large is the stronger model overall, scoring 31.9 to 29.0 on the Noometry Index. Granite 4.0 Micro costs 74× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Last verified . 7 shared benchmarks.

Granite 4.0 Micro IBM

29.0

Rank #318 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Granite 4.0 Micro scores higher in 3 categories and Mistral Large in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Large leads 30.1 to 9.9.
  • The biggest single-benchmark swing is GPQA Diamond: 28.3% for Granite 4.0 Micro and 51.3% for Mistral Large.
  • Granite 4.0 Micro is cheaper at $0.017 / $0.11 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Mistral Large accepts more context: 131K tokens versus 131K.

Side by side

Granite 4.0 Micro and Mistral Large specifications
Granite 4.0 MicroMistral Large
ProviderIBMMistral AI
Noometry Index29.031.9
Released2025-10-022024-02-26
WeightsOpenOpen
Context window131K131K
Max output118K16K
Input $ / M tokens$0.017$2
Output $ / M tokens$0.11$6
Results tracked851

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Granite 4.0 Micro: —, Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGranite 4.0 MicroMistral Large
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
LMArena Coding—1277
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Granite 4.0 Micro: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGranite 4.0 MicroMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Granite 4.0 Micro leads

Granite 4.0 Micro: 19.2 (#265), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGranite 4.0 MicroMistral Large
SimpleBench—22.5%
CritPt—0%
Chess Puzzles0%—
LiveBench Reasoning—43.5%
LMArena Hard Prompts—1257
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Mistral Large leads

Granite 4.0 Micro: 12.0 (#307), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGranite 4.0 MicroMistral Large
OTIS Mock AIME 2024-20252.8%8.5%
Omni-MATH20.9%28.1%
LiveBench Math—42.5%
LMArena Math—1262
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Mistral Large leads

Granite 4.0 Micro: 9.9 (#304), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGranite 4.0 MicroMistral Large
GPQA Diamond28.3%51.3%
MMLU-Pro39.5%59.9%
GPQA (HELM)30.7%43.5%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
LMArena Expert—1232
MMLU—80%

Multilingual Not comparable

Granite 4.0 Micro: —, Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGranite 4.0 MicroMistral Large
LMArena Non-English—1237
LMArena Chinese—1240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Granite 4.0 Micro leads

Granite 4.0 Micro: 69.9 (#169), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGranite 4.0 MicroMistral Large
IFEval84.9%87.7%
LiveBench Instruction Following—67.9%
LMArena Instruction Following—1249

Long Context Not comparable

Granite 4.0 Micro: —, Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGranite 4.0 MicroMistral Large
LMArena Longer Query—1261

Writing & Preference Granite 4.0 Micro leads

Granite 4.0 Micro: 46.7 (#216), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGranite 4.0 MicroMistral Large
WildBench67%80.1%
LMArena Text—1266
LMArena Creative Writing—1243
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
LMArena Multi-Turn—1260
LiveBench Language—39.4%

Frequently asked questions

Is Granite 4.0 Micro better than Mistral Large?

Mistral Large is the stronger model overall, scoring 31.9 to 29.0 on the Noometry Index. Granite 4.0 Micro costs 74× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Which is cheaper, Granite 4.0 Micro or Mistral Large?

Granite 4.0 Micro is cheaper. It lists at $0.017 per million input tokens and $0.11 per million output tokens; Mistral Large lists at $2 and $6.

Which has the bigger context window?

Mistral Large does, with 131K tokens against 131K.

How many benchmarks do Granite 4.0 Micro and Mistral Large share?

7 benchmarks have published results for both models. Granite 4.0 Micro has 8 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper