Model comparison

GPT-5-Codex vs Granite 4.1 8b

GPT-5-Codex and Granite 4.1 8b score almost the same on the Noometry Index (37.9 vs 37.4), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

GPT-5-Codex OpenAI

37.9

Rank #192 Reported

Granite 4.1 8b IBM

37.4

Rank #205 Confirmed

Summary

  • The widest gap is in coding, where GPT-5-Codex leads 42.4 to 30.1.
  • Granite 4.1 8b has downloadable open weights; the other is API-only.

Side by side

GPT-5-Codex and Granite 4.1 8b specifications
GPT-5-CodexGranite 4.1 8b
ProviderOpenAIIBM
Noometry Index37.937.4
Released2025-09-15—
WeightsProprietaryOpen
Context window400K—
Max output128K—
Input $ / M tokens$1.25—
Output $ / M tokens$10—
Results tracked313

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5-Codex leads

GPT-5-Codex: 42.4 (#103), Granite 4.1 8b: 30.1 (#297)

Coding benchmarks
BenchmarkGPT-5-CodexGranite 4.1 8b
LMArena WebDev—1192
WeirdML54.5%—
LMArena Coding—1312

Agentic & Tool Use Not comparable

GPT-5-Codex: 31.0 (#72), Granite 4.1 8b: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5-CodexGranite 4.1 8b
Terminal-Bench44.3%—

Reasoning GPT-5-Codex leads

GPT-5-Codex: 30.9 (#83), Granite 4.1 8b: 25.7 (#143)

Reasoning benchmarks
BenchmarkGPT-5-CodexGranite 4.1 8b
Kagi LLM Benchmark70.3%—
LMArena Hard Prompts—1293

Math Not comparable

GPT-5-Codex: —, Granite 4.1 8b: 36.4 (#166)

Math benchmarks
BenchmarkGPT-5-CodexGranite 4.1 8b
LMArena Math—1312

Knowledge Not comparable

GPT-5-Codex: —, Granite 4.1 8b: 36.1 (#174)

Knowledge benchmarks
BenchmarkGPT-5-CodexGranite 4.1 8b
LMArena Expert—1309

Multilingual Not comparable

GPT-5-Codex: —, Granite 4.1 8b: 41.7 (#204)

Multilingual benchmarks
BenchmarkGPT-5-CodexGranite 4.1 8b
LMArena Non-English—1261
LMArena Chinese—1337
LMArena Russian—1240

Instruction Following Not comparable

GPT-5-Codex: —, Granite 4.1 8b: 66.9 (#203)

Instruction Following benchmarks
BenchmarkGPT-5-CodexGranite 4.1 8b
LMArena Instruction Following—1269

Long Context Not comparable

GPT-5-Codex: —, Granite 4.1 8b: 38.7 (#193)

Long Context benchmarks
BenchmarkGPT-5-CodexGranite 4.1 8b
LMArena Longer Query—1275

Writing & Preference Not comparable

GPT-5-Codex: —, Granite 4.1 8b: 48.1 (#204)

Writing & Preference benchmarks
BenchmarkGPT-5-CodexGranite 4.1 8b
LMArena Text—1290
LMArena Creative Writing—1251
LMArena Multi-Turn—1268

Frequently asked questions

Is GPT-5-Codex better than Granite 4.1 8b?

GPT-5-Codex and Granite 4.1 8b score almost the same on the Noometry Index (37.9 vs 37.4), so choose on price, context window or the category you care about most.

Is GPT-5-Codex or Granite 4.1 8b better for coding?

GPT-5-Codex scores higher on coding benchmarks: 42.4 versus 30.1 in the Noometry coding category.

How many benchmarks do GPT-5-Codex and Granite 4.1 8b share?

0 benchmarks have published results for both models. GPT-5-Codex has 3 scored results on Noometry and Granite 4.1 8b has 13.

Related comparisons

Go deeper