Model comparison

Granite 4.2 30b vs Qwen Max

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 34.7 on the Noometry Index.

Last verified . 11 shared benchmarks.

Granite 4.2 30b IBM

41.8

Rank #130 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Granite 4.2 30b scores higher in 7 categories and Qwen Max in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Granite 4.2 30b leads 41.0 to 30.7.
  • Granite 4.2 30b has downloadable open weights; the other is API-only.

Side by side

Granite 4.2 30b and Qwen Max specifications
Granite 4.2 30bQwen Max
ProviderIBMAlibaba (Qwen)
Noometry Index41.834.7
Released—2024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked1123

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.2 30b leads

Granite 4.2 30b: 41.0 (#126), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkGranite 4.2 30bQwen Max
LMArena Coding13961288
Aider Polyglot—21.8%

Reasoning Granite 4.2 30b leads

Granite 4.2 30b: 27.8 (#112), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkGranite 4.2 30bQwen Max
LMArena Hard Prompts13741269

Math Not comparable

Granite 4.2 30b: —, Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkGranite 4.2 30bQwen Max
OTIS Mock AIME 2024-2025—16.1%
LMArena Math—1275
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Granite 4.2 30b leads

Granite 4.2 30b: 39.1 (#138), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkGranite 4.2 30bQwen Max
LMArena Expert14061248
GPQA Diamond—56.1%

Multilingual Granite 4.2 30b leads

Granite 4.2 30b: 47.3 (#151), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkGranite 4.2 30bQwen Max
LMArena Non-English13401263
LMArena Chinese14141254
LMArena Russian13431274
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Spanish—1290

Instruction Following Granite 4.2 30b leads

Granite 4.2 30b: 71.2 (#155), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkGranite 4.2 30bQwen Max
LMArena Instruction Following13471262

Long Context Granite 4.2 30b leads

Granite 4.2 30b: 41.4 (#140), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkGranite 4.2 30bQwen Max
LMArena Longer Query13591288
Fiction.LiveBench—66.7%

Writing & Preference Granite 4.2 30b leads

Granite 4.2 30b: 53.8 (#156), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkGranite 4.2 30bQwen Max
LMArena Text13611282
LMArena Creative Writing12881248
LMArena Multi-Turn13391277

Frequently asked questions

Is Granite 4.2 30b better than Qwen Max?

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 34.7 on the Noometry Index.

Is Granite 4.2 30b or Qwen Max better for coding?

Granite 4.2 30b scores higher on coding benchmarks: 41.0 versus 30.7 in the Noometry coding category.

How many benchmarks do Granite 4.2 30b and Qwen Max share?

11 benchmarks have published results for both models. Granite 4.2 30b has 11 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper