Model comparison

Granite 4.2 30b vs Qwen2.5 Plus 1127

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 38.8 on the Noometry Index.

Last verified . 11 shared benchmarks.

Granite 4.2 30b IBM

41.8

Rank #130 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Granite 4.2 30b scores higher in 7 categories and Qwen2.5 Plus 1127 in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Granite 4.2 30b leads 47.3 to 41.9.
  • Granite 4.2 30b has downloadable open weights; the other is API-only.

Side by side

Granite 4.2 30b and Qwen2.5 Plus 1127 specifications
Granite 4.2 30bQwen2.5 Plus 1127
ProviderIBMAlibaba (Qwen)
Noometry Index41.838.8
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1114

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.2 30b leads

Granite 4.2 30b: 41.0 (#126), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkGranite 4.2 30bQwen2.5 Plus 1127
LMArena Coding13961314

Reasoning Granite 4.2 30b leads

Granite 4.2 30b: 27.8 (#112), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkGranite 4.2 30bQwen2.5 Plus 1127
LMArena Hard Prompts13741299

Math Not comparable

Granite 4.2 30b: —, Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkGranite 4.2 30bQwen2.5 Plus 1127
LMArena Math—1298

Knowledge Granite 4.2 30b leads

Granite 4.2 30b: 39.1 (#138), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkGranite 4.2 30bQwen2.5 Plus 1127
LMArena Expert14061289

Multilingual Granite 4.2 30b leads

Granite 4.2 30b: 47.3 (#151), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkGranite 4.2 30bQwen2.5 Plus 1127
LMArena Non-English13401265
LMArena Chinese14141314
LMArena Russian13431271
LMArena German—1231
LMArena Japanese—1207

Instruction Following Granite 4.2 30b leads

Granite 4.2 30b: 71.2 (#155), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkGranite 4.2 30bQwen2.5 Plus 1127
LMArena Instruction Following13471275

Long Context Granite 4.2 30b leads

Granite 4.2 30b: 41.4 (#140), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkGranite 4.2 30bQwen2.5 Plus 1127
LMArena Longer Query13591292

Writing & Preference Granite 4.2 30b leads

Granite 4.2 30b: 53.8 (#156), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkGranite 4.2 30bQwen2.5 Plus 1127
LMArena Text13611299
LMArena Creative Writing12881262
LMArena Multi-Turn13391299

Frequently asked questions

Is Granite 4.2 30b better than Qwen2.5 Plus 1127?

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 38.8 on the Noometry Index.

Is Granite 4.2 30b or Qwen2.5 Plus 1127 better for coding?

Granite 4.2 30b scores higher on coding benchmarks: 41.0 versus 38.5 in the Noometry coding category.

How many benchmarks do Granite 4.2 30b and Qwen2.5 Plus 1127 share?

11 benchmarks have published results for both models. Granite 4.2 30b has 11 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper