Model comparison

GPT-5.1-Codex vs Granite 4.2 3b

GPT-5.1-Codex and Granite 4.2 3b score almost the same on the Noometry Index (38.6 vs 39.4), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

GPT-5.1-Codex OpenAI

38.6

Rank #186 Reported

Granite 4.2 3b IBM

39.4

Rank #169 Confirmed

Summary

  • Granite 4.2 3b has downloadable open weights; the other is API-only.

Side by side

GPT-5.1-Codex and Granite 4.2 3b specifications
GPT-5.1-CodexGranite 4.2 3b
ProviderOpenAIIBM
Noometry Index38.639.4
Released2025-11-12—
WeightsProprietaryOpen
Context window400K—
Max output128K—
Input $ / M tokens$1.25—
Output $ / M tokens$10—
Results tracked611

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.1-Codex leads

GPT-5.1-Codex: 41.9 (#116), Granite 4.2 3b: 40.0 (#151)

Coding benchmarks
BenchmarkGPT-5.1-CodexGranite 4.2 3b
SWE-bench Verified (bash only)66%—
LMArena WebDev1337—
LMArena Coding—1361
ALE-Bench1,245—

Agentic & Tool Use Not comparable

GPT-5.1-Codex: 38.0 (#33), Granite 4.2 3b: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.1-CodexGranite 4.2 3b
Terminal-Bench60.4%—
METR Time Horizons70.8%—

Reasoning Not comparable

GPT-5.1-Codex: —, Granite 4.2 3b: 26.0 (#138)

Reasoning benchmarks
BenchmarkGPT-5.1-CodexGranite 4.2 3b
LMArena Hard Prompts—1306

Math Not comparable

GPT-5.1-Codex: 30.3 (#235), Granite 4.2 3b: —

Math benchmarks
BenchmarkGPT-5.1-CodexGranite 4.2 3b
ProofBench9%—

Knowledge Not comparable

GPT-5.1-Codex: —, Granite 4.2 3b: 36.3 (#171)

Knowledge benchmarks
BenchmarkGPT-5.1-CodexGranite 4.2 3b
LMArena Expert—1315

Multilingual Not comparable

GPT-5.1-Codex: —, Granite 4.2 3b: 42.1 (#198)

Multilingual benchmarks
BenchmarkGPT-5.1-CodexGranite 4.2 3b
LMArena Non-English—1268
LMArena Chinese—1269
LMArena Russian—1249

Instruction Following Not comparable

GPT-5.1-Codex: —, Granite 4.2 3b: 67.1 (#200)

Instruction Following benchmarks
BenchmarkGPT-5.1-CodexGranite 4.2 3b
LMArena Instruction Following—1273

Long Context Not comparable

GPT-5.1-Codex: —, Granite 4.2 3b: 39.2 (#185)

Long Context benchmarks
BenchmarkGPT-5.1-CodexGranite 4.2 3b
LMArena Longer Query—1291

Writing & Preference Not comparable

GPT-5.1-Codex: —, Granite 4.2 3b: 47.2 (#212)

Writing & Preference benchmarks
BenchmarkGPT-5.1-CodexGranite 4.2 3b
LMArena Text—1293
LMArena Creative Writing—1205
LMArena Multi-Turn—1290

Frequently asked questions

Is GPT-5.1-Codex better than Granite 4.2 3b?

GPT-5.1-Codex and Granite 4.2 3b score almost the same on the Noometry Index (38.6 vs 39.4), so choose on price, context window or the category you care about most.

Is GPT-5.1-Codex or Granite 4.2 3b better for coding?

GPT-5.1-Codex scores higher on coding benchmarks: 41.9 versus 40.0 in the Noometry coding category.

How many benchmarks do GPT-5.1-Codex and Granite 4.2 3b share?

0 benchmarks have published results for both models. GPT-5.1-Codex has 6 scored results on Noometry and Granite 4.2 3b has 11.

Related comparisons

Go deeper