Model comparison

Granite 4.2 8B vs Step 3.7 Flash

Granite 4.2 8B is the stronger model overall, scoring 40.5 to 37.3 on the Noometry Index.

Last verified . 0 shared benchmarks.

Granite 4.2 8B IBM

40.5

Rank #148 Confirmed

Step 3.7 Flash StepFun

37.3

Rank #207 Reported

Summary

  • The widest gap is in reasoning, where Granite 4.2 8B leads 26.6 to 21.6.
  • Granite 4.2 8B is cheaper at $0.06 / $0.25 per million input/output tokens, against $0.18 / $1.11 for Step 3.7 Flash.
  • Step 3.7 Flash accepts more context: 256K tokens versus 131K.

Side by side

Granite 4.2 8B and Step 3.7 Flash specifications
Granite 4.2 8BStep 3.7 Flash
ProviderIBMStepFun
Noometry Index40.537.3
Released—2026-05-29
WeightsOpenOpen
Context window131K256K
Max output118K256K
Input $ / M tokens$0.06$0.18
Output $ / M tokens$0.25$1.11
Results tracked115

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Granite 4.2 8B: 40.5 (#137), Step 3.7 Flash: 40.0 (#150)

Coding benchmarks
BenchmarkGranite 4.2 8BStep 3.7 Flash
SciCode—40%
LMArena Coding1380—
ALE-Bench—694.12

Reasoning Granite 4.2 8B leads

Granite 4.2 8B: 26.6 (#131), Step 3.7 Flash: 21.6 (#219)

Reasoning benchmarks
BenchmarkGranite 4.2 8BStep 3.7 Flash
NYT Connections (extended)—39.7%
CritPt—2.3%
LMArena Hard Prompts1329—

Math Not comparable

Granite 4.2 8B: —, Step 3.7 Flash: 42.9 (#82)

Math benchmarks
BenchmarkGranite 4.2 8BStep 3.7 Flash
MathArena Final-Answer Competitions—68.5%

Knowledge Not comparable

Granite 4.2 8B: 38.4 (#145), Step 3.7 Flash: —

Knowledge benchmarks
BenchmarkGranite 4.2 8BStep 3.7 Flash
LMArena Expert1384—

Multilingual Not comparable

Granite 4.2 8B: 44.5 (#178), Step 3.7 Flash: —

Multilingual benchmarks
BenchmarkGranite 4.2 8BStep 3.7 Flash
LMArena Non-English1302—
LMArena Chinese1366—
LMArena Russian1285—

Instruction Following Not comparable

Granite 4.2 8B: 68.7 (#184), Step 3.7 Flash: —

Instruction Following benchmarks
BenchmarkGranite 4.2 8BStep 3.7 Flash
LMArena Instruction Following1301—

Long Context Not comparable

Granite 4.2 8B: 40.3 (#159), Step 3.7 Flash: —

Long Context benchmarks
BenchmarkGranite 4.2 8BStep 3.7 Flash
LMArena Longer Query1324—

Writing & Preference Not comparable

Granite 4.2 8B: 49.6 (#189), Step 3.7 Flash: —

Writing & Preference benchmarks
BenchmarkGranite 4.2 8BStep 3.7 Flash
LMArena Text1320—
LMArena Creative Writing1236—
LMArena Multi-Turn1301—

Frequently asked questions

Is Granite 4.2 8B better than Step 3.7 Flash?

Granite 4.2 8B is the stronger model overall, scoring 40.5 to 37.3 on the Noometry Index.

Which is cheaper, Granite 4.2 8B or Step 3.7 Flash?

Granite 4.2 8B is cheaper. It lists at $0.06 per million input tokens and $0.25 per million output tokens; Step 3.7 Flash lists at $0.18 and $1.11.

Is Granite 4.2 8B or Step 3.7 Flash better for coding?

They score almost the same on coding (40.5 vs 40.0); test both on your own repository before choosing.

Which has the bigger context window?

Step 3.7 Flash does, with 256K tokens against 131K.

How many benchmarks do Granite 4.2 8B and Step 3.7 Flash share?

0 benchmarks have published results for both models. Granite 4.2 8B has 11 scored results on Noometry and Step 3.7 Flash has 5.

Related comparisons

Go deeper