Model comparison

Granite 4.2 30b vs Step 3

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 40.5 on the Noometry Index.

Last verified . 11 shared benchmarks.

Granite 4.2 30b IBM

41.8

Rank #130 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Granite 4.2 30b scores higher in 5 categories and Step 3 in 2 categories; 2 gaps are clear of the uncertainty.

Side by side

Granite 4.2 30b and Step 3 specifications
Granite 4.2 30bStep 3
ProviderIBMStepFun
Noometry Index41.840.5
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1117

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Granite 4.2 30b: 41.0 (#126), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkGranite 4.2 30bStep 3
LMArena Coding13961367

Reasoning Too close to call

Granite 4.2 30b: 27.8 (#112), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkGranite 4.2 30bStep 3
LMArena Hard Prompts13741355
Kagi LLM Benchmark—62.3%

Math Not comparable

Granite 4.2 30b: —, Step 3: 37.6 (#148)

Math benchmarks
BenchmarkGranite 4.2 30bStep 3
LMArena Math—1366

Knowledge Granite 4.2 30b leads

Granite 4.2 30b: 39.1 (#138), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkGranite 4.2 30bStep 3
LMArena Expert14061333

Multimodal Not comparable

Granite 4.2 30b: —, Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkGranite 4.2 30bStep 3
LMArena Vision—1177

Multilingual Too close to call

Granite 4.2 30b: 47.3 (#151), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkGranite 4.2 30bStep 3
LMArena Non-English13401327
LMArena Chinese14141397
LMArena Russian13431331
LMArena German—1371
LMArena Korean—1269
LMArena Spanish—1371

Instruction Following Too close to call

Granite 4.2 30b: 71.2 (#155), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkGranite 4.2 30bStep 3
LMArena Instruction Following13471332

Long Context Granite 4.2 30b leads

Granite 4.2 30b: 41.4 (#140), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkGranite 4.2 30bStep 3
LMArena Longer Query13591326

Writing & Preference Too close to call

Granite 4.2 30b: 53.8 (#156), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkGranite 4.2 30bStep 3
LMArena Text13611350
LMArena Creative Writing12881321
LMArena Multi-Turn13391341

Frequently asked questions

Is Granite 4.2 30b better than Step 3?

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 40.5 on the Noometry Index.

Is Granite 4.2 30b or Step 3 better for coding?

They score almost the same on coding (41.0 vs 40.1); test both on your own repository before choosing.

How many benchmarks do Granite 4.2 30b and Step 3 share?

11 benchmarks have published results for both models. Granite 4.2 30b has 11 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper