Model comparison

Granite 4.2 30b vs o3

o3 is the stronger model overall, scoring 47.5 to 41.8 on the Noometry Index.

Last verified . 11 shared benchmarks.

Granite 4.2 30b IBM

41.8

Rank #130 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Granite 4.2 30b scores higher in 0 categories and o3 in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 39.1.
  • Granite 4.2 30b has downloadable open weights; the other is API-only.

Side by side

Granite 4.2 30b and o3 specifications
Granite 4.2 30bo3
ProviderIBMOpenAI
Noometry Index41.847.5
Released—2025-04-16
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked1163

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Granite 4.2 30b: 41.0 (#126), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGranite 4.2 30bo3
LMArena Coding13961408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Granite 4.2 30b: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGranite 4.2 30bo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Granite 4.2 30b: 27.8 (#112), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGranite 4.2 30bo3
LMArena Hard Prompts13741402
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math Not comparable

Granite 4.2 30b: —, o3: 50.2 (#58)

Math benchmarks
BenchmarkGranite 4.2 30bo3
FrontierMath (Tiers 1-3)—33.3%
OTIS Mock AIME 2024-2025—84.4%
Omni-MATH—71.4%
LMArena Math—1426
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Granite 4.2 30b: 39.1 (#138), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGranite 4.2 30bo3
LMArena Expert14061402
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%

Multimodal Not comparable

Granite 4.2 30b: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGranite 4.2 30bo3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual o3 leads

Granite 4.2 30b: 47.3 (#151), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGranite 4.2 30bo3
LMArena Non-English13401401
LMArena Chinese14141437
LMArena Russian13431406
LMArena French—1430
LMArena German—1420
LMArena Japanese—1403
LMArena Korean—1370
LMArena Spanish—1395

Instruction Following o3 leads

Granite 4.2 30b: 71.2 (#155), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGranite 4.2 30bo3
LMArena Instruction Following13471368
IFEval—86.9%

Long Context o3 leads

Granite 4.2 30b: 41.4 (#140), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGranite 4.2 30bo3
LMArena Longer Query13591372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Granite 4.2 30b: 53.8 (#156), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGranite 4.2 30bo3
LMArena Text13611410
LMArena Creative Writing12881359
LMArena Multi-Turn13391405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Granite 4.2 30b better than o3?

o3 is the stronger model overall, scoring 47.5 to 41.8 on the Noometry Index.

Is Granite 4.2 30b or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 41.0 in the Noometry coding category.

How many benchmarks do Granite 4.2 30b and o3 share?

11 benchmarks have published results for both models. Granite 4.2 30b has 11 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper