Model comparison

Granite 4.0 H Small vs Mistral Medium

Granite 4.0 H Small and Mistral Medium score almost the same on the Noometry Index (36.5 vs 36.3), so choose on price, context window or the category you care about most.

Last verified . 14 shared benchmarks.

Granite 4.0 H Small IBM

36.5

Rank #214 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Granite 4.0 H Small scores higher in 4 categories and Mistral Medium in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 43.8.
  • The biggest single-benchmark swing is Vectara Hallucination Rate: 5.2% for Granite 4.0 H Small and 22.7% for Mistral Medium.

Side by side

Granite 4.0 H Small and Mistral Medium specifications
Granite 4.0 H SmallMistral Medium
ProviderIBMMistral AI
Noometry Index36.536.3
Released—2023-12-11
WeightsOpenOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked1936

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.0 H Small leads

Granite 4.0 H Small: 36.4 (#209), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkGranite 4.0 H SmallMistral Medium
LMArena Coding12491434
FrontierCode—8%
SciCode—40.2%
WeirdML—43.7%
ALE-Bench—763.98

Agentic & Tool Use Not comparable

Granite 4.0 H Small: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkGranite 4.0 H SmallMistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning Too close to call

Granite 4.0 H Small: 24.4 (#163), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkGranite 4.0 H SmallMistral Medium
LMArena Hard Prompts12401426
Kagi LLM Benchmark—50%
CritPt—0%
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Granite 4.0 H Small leads

Granite 4.0 H Small: 31.4 (#223), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkGranite 4.0 H SmallMistral Medium
LMArena Math12471408
OTIS Mock AIME 2024-2025—32.2%
ProofBench—9%
Omni-MATH29.6%—
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Granite 4.0 H Small leads

Granite 4.0 H Small: 32.0 (#215), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkGranite 4.0 H SmallMistral Medium
Vectara Hallucination Rate5.2%22.7%
LMArena Expert12511408
GPQA Diamond—59.5%
Humanity's Last Exam—4.5%
MMLU-Pro56.9%—
GPQA (HELM)38.3%—

Multimodal Not comparable

Granite 4.0 H Small: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkGranite 4.0 H SmallMistral Medium
LMArena Vision—1172

Multilingual Mistral Medium leads

Granite 4.0 H Small: 38.6 (#228), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkGranite 4.0 H SmallMistral Medium
LMArena Non-English12161408
LMArena Chinese12491447
LMArena Russian12011411
LMArena Spanish12581433
LMArena French—1459
LMArena German—1432
LMArena Japanese—1378
LMArena Korean—1380

Instruction Following Mistral Medium leads

Granite 4.0 H Small: 68.7 (#183), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkGranite 4.0 H SmallMistral Medium
LMArena Instruction Following12221398
IFEval89%—

Long Context Mistral Medium leads

Granite 4.0 H Small: 37.7 (#213), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkGranite 4.0 H SmallMistral Medium
LMArena Longer Query12421406

Writing & Preference Mistral Medium leads

Granite 4.0 H Small: 43.8 (#227), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkGranite 4.0 H SmallMistral Medium
LMArena Text12411424
LMArena Creative Writing12111391
LMArena Multi-Turn12421418
Short-Story Creative Writing—77.3%
WildBench73.9%—

Frequently asked questions

Is Granite 4.0 H Small better than Mistral Medium?

Granite 4.0 H Small and Mistral Medium score almost the same on the Noometry Index (36.5 vs 36.3), so choose on price, context window or the category you care about most.

Is Granite 4.0 H Small or Mistral Medium better for coding?

Granite 4.0 H Small scores higher on coding benchmarks: 36.4 versus 34.2 in the Noometry coding category.

How many benchmarks do Granite 4.0 H Small and Mistral Medium share?

14 benchmarks have published results for both models. Granite 4.0 H Small has 19 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper