Model comparison

Granite 4.0 Micro vs Mistral Large 3

Mistral Large 3 is the stronger model overall, scoring 39.1 to 29.0 on the Noometry Index. Granite 4.0 Micro costs 9.2× less per token, which makes it the better buy when Mistral Large 3's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Granite 4.0 Micro IBM

29.0

Rank #318 Confirmed

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Summary

  • The widest gap is in math, where Mistral Large 3 leads 38.7 to 12.0.
  • Granite 4.0 Micro is cheaper at $0.017 / $0.11 per million input/output tokens, against $0.25 / $0.75 for Mistral Large 3.
  • Mistral Large 3 accepts more context: 262K tokens versus 131K.

Side by side

Granite 4.0 Micro and Mistral Large 3 specifications
Granite 4.0 MicroMistral Large 3
ProviderIBMMistral AI
Noometry Index29.039.1
Released2025-10-022025-12-02
WeightsOpenOpen
Context window131K262K
Max output118K8K
Input $ / M tokens$0.017$0.25
Output $ / M tokens$0.11$0.75
Results tracked824

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Granite 4.0 Micro: —, Mistral Large 3: 34.4 (#237)

Coding benchmarks
BenchmarkGranite 4.0 MicroMistral Large 3
LMArena WebDev—1230
LMArena Coding—1448

Reasoning Granite 4.0 Micro leads

Granite 4.0 Micro: 19.2 (#265), Mistral Large 3: 15.2 (#319)

Reasoning benchmarks
BenchmarkGranite 4.0 MicroMistral Large 3
Kagi LLM Benchmark—50.9%
NYT Connections (extended)—7.5%
Chess Puzzles0%—
Thematic Generalization—23%
LMArena Hard Prompts—1429

Math Mistral Large 3 leads

Granite 4.0 Micro: 12.0 (#307), Mistral Large 3: 38.7 (#129)

Math benchmarks
BenchmarkGranite 4.0 MicroMistral Large 3
OTIS Mock AIME 2024-20252.8%—
Omni-MATH20.9%—
LMArena Math—1414

Knowledge Mistral Large 3 leads

Granite 4.0 Micro: 9.9 (#304), Mistral Large 3: 36.0 (#177)

Knowledge benchmarks
BenchmarkGranite 4.0 MicroMistral Large 3
GPQA Diamond28.3%—
MMLU-Pro39.5%—
Vectara Hallucination Rate—14.5%
GPQA (HELM)30.7%—
LMArena Expert—1421

Multimodal Not comparable

Granite 4.0 Micro: —, Mistral Large 3: 38.2 (#66)

Multimodal benchmarks
BenchmarkGranite 4.0 MicroMistral Large 3
LMArena Vision—1221

Multilingual Not comparable

Granite 4.0 Micro: —, Mistral Large 3: 52.5 (#84)

Multilingual benchmarks
BenchmarkGranite 4.0 MicroMistral Large 3
LMArena Non-English—1413
LMArena Chinese—1447
LMArena French—1455
LMArena German—1437
LMArena Japanese—1394
LMArena Korean—1384
LMArena Russian—1411
LMArena Spanish—1440

Instruction Following Mistral Large 3 leads

Granite 4.0 Micro: 69.9 (#169), Mistral Large 3: 74.0 (#108)

Instruction Following benchmarks
BenchmarkGranite 4.0 MicroMistral Large 3
IFEval84.9%—
LMArena Instruction Following—1403

Long Context Not comparable

Granite 4.0 Micro: —, Mistral Large 3: 43.1 (#105)

Long Context benchmarks
BenchmarkGranite 4.0 MicroMistral Large 3
LMArena Longer Query—1413

Writing & Preference Mistral Large 3 leads

Granite 4.0 Micro: 46.7 (#216), Mistral Large 3: 60.0 (#101)

Writing & Preference benchmarks
BenchmarkGranite 4.0 MicroMistral Large 3
LMArena Text—1428
LMArena Creative Writing—1386
EQ-Bench Creative Writing—1412
WildBench67%—
LMArena Multi-Turn—1429

Frequently asked questions

Is Granite 4.0 Micro better than Mistral Large 3?

Mistral Large 3 is the stronger model overall, scoring 39.1 to 29.0 on the Noometry Index. Granite 4.0 Micro costs 9.2× less per token, which makes it the better buy when Mistral Large 3's lead doesn't matter for your workload.

Which is cheaper, Granite 4.0 Micro or Mistral Large 3?

Granite 4.0 Micro is cheaper. It lists at $0.017 per million input tokens and $0.11 per million output tokens; Mistral Large 3 lists at $0.25 and $0.75.

Which has the bigger context window?

Mistral Large 3 does, with 262K tokens against 131K.

How many benchmarks do Granite 4.0 Micro and Mistral Large 3 share?

0 benchmarks have published results for both models. Granite 4.0 Micro has 8 scored results on Noometry and Mistral Large 3 has 24.

Related comparisons

Go deeper