Model comparison

Granite 4.0 H Small vs Mistral Large

Granite 4.0 H Small is the stronger model overall, scoring 36.5 to 31.9 on the Noometry Index.

Last verified . 19 shared benchmarks.

Granite 4.0 H Small IBM

36.5

Rank #214 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Granite 4.0 H Small scores higher in 6 categories and Mistral Large in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Granite 4.0 H Small leads 31.4 to 18.2.
  • The biggest single-benchmark swing is WildBench: 73.9% for Granite 4.0 H Small and 80.1% for Mistral Large.

Side by side

Granite 4.0 H Small and Mistral Large specifications
Granite 4.0 H SmallMistral Large
ProviderIBMMistral AI
Noometry Index36.531.9
Released—2024-02-26
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1951

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.0 H Small leads

Granite 4.0 H Small: 36.4 (#209), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGranite 4.0 H SmallMistral Large
LMArena Coding12491277
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Granite 4.0 H Small: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGranite 4.0 H SmallMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Granite 4.0 H Small leads

Granite 4.0 H Small: 24.4 (#163), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGranite 4.0 H SmallMistral Large
LMArena Hard Prompts12401257
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Granite 4.0 H Small leads

Granite 4.0 H Small: 31.4 (#223), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGranite 4.0 H SmallMistral Large
Omni-MATH29.6%28.1%
LMArena Math12471262
OTIS Mock AIME 2024-2025—8.5%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Granite 4.0 H Small leads

Granite 4.0 H Small: 32.0 (#215), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGranite 4.0 H SmallMistral Large
MMLU-Pro56.9%59.9%
Vectara Hallucination Rate5.2%4.5%
GPQA (HELM)38.3%43.5%
LMArena Expert12511232
GPQA Diamond—51.3%
Confabulations—21.4%
MMLU—80%

Multilingual Mistral Large leads

Granite 4.0 H Small: 38.6 (#228), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGranite 4.0 H SmallMistral Large
LMArena Non-English12161237
LMArena Chinese12491240
LMArena Russian12011257
LMArena Spanish12581268
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202

Instruction Following Too close to call

Granite 4.0 H Small: 68.7 (#183), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGranite 4.0 H SmallMistral Large
IFEval89%87.7%
LMArena Instruction Following12221249
LiveBench Instruction Following—67.9%

Long Context Too close to call

Granite 4.0 H Small: 37.7 (#213), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGranite 4.0 H SmallMistral Large
LMArena Longer Query12421261

Writing & Preference Granite 4.0 H Small leads

Granite 4.0 H Small: 43.8 (#227), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGranite 4.0 H SmallMistral Large
LMArena Text12411266
LMArena Creative Writing12111243
WildBench73.9%80.1%
LMArena Multi-Turn12421260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
LiveBench Language—39.4%

Frequently asked questions

Is Granite 4.0 H Small better than Mistral Large?

Granite 4.0 H Small is the stronger model overall, scoring 36.5 to 31.9 on the Noometry Index.

Is Granite 4.0 H Small or Mistral Large better for coding?

Granite 4.0 H Small scores higher on coding benchmarks: 36.4 versus 34.3 in the Noometry coding category.

How many benchmarks do Granite 4.0 H Small and Mistral Large share?

19 benchmarks have published results for both models. Granite 4.0 H Small has 19 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper