Model comparison

GPT-4 Turbo vs Granite 4.0 H Small

Granite 4.0 H Small is the stronger model overall, scoring 36.5 to 30.5 on the Noometry Index.

Last verified . 13 shared benchmarks.

GPT-4 Turbo OpenAI

30.5

Rank #292 Confirmed

Granite 4.0 H Small IBM

36.5

Rank #214 Confirmed

Summary

  • They share 13 benchmarks with published results for both. GPT-4 Turbo scores higher in 3 categories and Granite 4.0 H Small in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Granite 4.0 H Small leads 31.4 to 9.0.
  • Granite 4.0 H Small has downloadable open weights; the other is API-only.

Side by side

GPT-4 Turbo and Granite 4.0 H Small specifications
GPT-4 TurboGranite 4.0 H Small
ProviderOpenAIIBM
Noometry Index30.536.5
Released2023-11-06—
WeightsProprietaryOpen
Context window128K—
Max output4K—
Input $ / M tokens$10—
Output $ / M tokens$30—
Results tracked3619

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.0 H Small leads

GPT-4 Turbo: 33.8 (#249), Granite 4.0 H Small: 36.4 (#209)

Coding benchmarks
BenchmarkGPT-4 TurboGranite 4.0 H Small
LMArena Coding12681249
WeirdML18%—
BigCodeBench Instruct48.2%—
BigCodeBench Complete58.2%—
HumanEval+86.6%—
MBPP+73.3%—

Agentic & Tool Use Not comparable

GPT-4 Turbo: —, Granite 4.0 H Small: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4 TurboGranite 4.0 H Small
METR Time Horizons36.7%—

Reasoning Granite 4.0 H Small leads

GPT-4 Turbo: 15.3 (#317), Granite 4.0 H Small: 24.4 (#163)

Reasoning benchmarks
BenchmarkGPT-4 TurboGranite 4.0 H Small
LMArena Hard Prompts12511240
SimpleBench25.1%—
Chess Puzzles6%—
DTBench61.6%—
LMCA9.8%—
Epoch Capabilities Index127.25—
ForecastBench59.4—

Math Granite 4.0 H Small leads

GPT-4 Turbo: 9.0 (#322), Granite 4.0 H Small: 31.4 (#223)

Math benchmarks
BenchmarkGPT-4 TurboGranite 4.0 H Small
LMArena Math12721247
FrontierMath (Tiers 1-3)0.7%—
OTIS Mock AIME 2024-20256.7%—
Omni-MATH—29.6%
MATH Level 546.7%—

Knowledge Granite 4.0 H Small leads

GPT-4 Turbo: 24.3 (#268), Granite 4.0 H Small: 32.0 (#215)

Knowledge benchmarks
BenchmarkGPT-4 TurboGranite 4.0 H Small
LMArena Expert12231251
GPQA Diamond46.6%—
MMLU-Pro—56.9%
Confabulations28.4%—
Vectara Hallucination Rate—5.2%
GPQA (HELM)—38.3%
MMLU81.3%—

Multimodal Not comparable

GPT-4 Turbo: 30.6 (#110), Granite 4.0 H Small: —

Multimodal benchmarks
BenchmarkGPT-4 TurboGranite 4.0 H Small
LMArena Vision1090—

Multilingual GPT-4 Turbo leads

GPT-4 Turbo: 40.5 (#216), Granite 4.0 H Small: 38.6 (#228)

Multilingual benchmarks
BenchmarkGPT-4 TurboGranite 4.0 H Small
LMArena Non-English12451216
LMArena Chinese12421249
LMArena Russian12591201
LMArena Spanish12601258
LMArena French1276—
LMArena German1259—
LMArena Japanese1194—
LMArena Korean1187—

Instruction Following Granite 4.0 H Small leads

GPT-4 Turbo: 65.8 (#216), Granite 4.0 H Small: 68.7 (#183)

Instruction Following benchmarks
BenchmarkGPT-4 TurboGranite 4.0 H Small
LMArena Instruction Following12491222
IFEval—89%

Long Context Too close to call

GPT-4 Turbo: 38.0 (#206), Granite 4.0 H Small: 37.7 (#213)

Long Context benchmarks
BenchmarkGPT-4 TurboGranite 4.0 H Small
LMArena Longer Query12541242

Writing & Preference GPT-4 Turbo leads

GPT-4 Turbo: 47.7 (#206), Granite 4.0 H Small: 43.8 (#227)

Writing & Preference benchmarks
BenchmarkGPT-4 TurboGranite 4.0 H Small
LMArena Text12721241
LMArena Creative Writing12691211
LMArena Multi-Turn12671242
WildBench—73.9%

Frequently asked questions

Is GPT-4 Turbo better than Granite 4.0 H Small?

Granite 4.0 H Small is the stronger model overall, scoring 36.5 to 30.5 on the Noometry Index.

Is GPT-4 Turbo or Granite 4.0 H Small better for coding?

Granite 4.0 H Small scores higher on coding benchmarks: 36.4 versus 33.8 in the Noometry coding category.

How many benchmarks do GPT-4 Turbo and Granite 4.0 H Small share?

13 benchmarks have published results for both models. GPT-4 Turbo has 36 scored results on Noometry and Granite 4.0 H Small has 19.

Related comparisons

Go deeper