Model comparison

Granite 4.2 30b vs Qwen3.6 Plus

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 41.8 on the Noometry Index.

Last verified . 11 shared benchmarks.

Granite 4.2 30b IBM

41.8

Rank #130 Confirmed

Qwen3.6 Plus Alibaba (Qwen)

47.5

Rank #62 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Granite 4.2 30b scores higher in 1 category and Qwen3.6 Plus in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.6 Plus leads 56.1 to 39.1.
  • Granite 4.2 30b has downloadable open weights; the other is API-only.

Side by side

Granite 4.2 30b and Qwen3.6 Plus specifications
Granite 4.2 30bQwen3.6 Plus
ProviderIBMAlibaba (Qwen)
Noometry Index41.847.5
Released—2026-03-31
WeightsOpenProprietary
Context window—1M
Max output—66K
Input $ / M tokens—$0.50
Output $ / M tokens—$3
Results tracked1137

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Granite 4.2 30b: 41.0 (#126), Qwen3.6 Plus: 40.8 (#130)

Coding benchmarks
BenchmarkGranite 4.2 30bQwen3.6 Plus
LMArena Coding13961467
SWE-bench Verified—57.9%
LMArena WebDev—1461
SciCode—40.7%
ALE-Bench—670.15

Agentic & Tool Use Not comparable

Granite 4.2 30b: —, Qwen3.6 Plus: —

Agentic & Tool Use benchmarks
BenchmarkGranite 4.2 30bQwen3.6 Plus
Vending-Bench 2—5,115

Reasoning Qwen3.6 Plus leads

Granite 4.2 30b: 27.8 (#112), Qwen3.6 Plus: 29.3 (#93)

Reasoning benchmarks
BenchmarkGranite 4.2 30bQwen3.6 Plus
LMArena Hard Prompts13741449
NYT Connections (extended)—60.3%
CritPt—2.9%
Chess Puzzles—17%
Thematic Generalization—59.5%
Mystery Game Puzzles—12%
DTBench—81.9%
LMCA—33.1%
Epoch Capabilities Index—147.65

Math Not comparable

Granite 4.2 30b: —, Qwen3.6 Plus: 51.8 (#54)

Math benchmarks
BenchmarkGranite 4.2 30bQwen3.6 Plus
FrontierMath (Tiers 1-3)—38.2%
OTIS Mock AIME 2024-2025—93.3%
LMArena Math—1450
FrontierMath (Feb 2025 set)—26.2%
FrontierMath Tier 4 (v1)—8.3%

Knowledge Qwen3.6 Plus leads

Granite 4.2 30b: 39.1 (#138), Qwen3.6 Plus: 56.1 (#45)

Knowledge benchmarks
BenchmarkGranite 4.2 30bQwen3.6 Plus
LMArena Expert14061454
GPQA Diamond—88.4%
SimpleQA Verified—44.1%

Multilingual Qwen3.6 Plus leads

Granite 4.2 30b: 47.3 (#151), Qwen3.6 Plus: 53.3 (#70)

Multilingual benchmarks
BenchmarkGranite 4.2 30bQwen3.6 Plus
LMArena Non-English13401424
LMArena Chinese14141477
LMArena Russian13431434
LMArena French—1455
LMArena German—1452
LMArena Japanese—1389
LMArena Korean—1379
LMArena Spanish—1432

Instruction Following Qwen3.6 Plus leads

Granite 4.2 30b: 71.2 (#155), Qwen3.6 Plus: 75.0 (#74)

Instruction Following benchmarks
BenchmarkGranite 4.2 30bQwen3.6 Plus
LMArena Instruction Following13471425

Long Context Qwen3.6 Plus leads

Granite 4.2 30b: 41.4 (#140), Qwen3.6 Plus: 45.2 (#49)

Long Context benchmarks
BenchmarkGranite 4.2 30bQwen3.6 Plus
LMArena Longer Query13591439
CL-bench—20.3%

Writing & Preference Qwen3.6 Plus leads

Granite 4.2 30b: 53.8 (#156), Qwen3.6 Plus: 62.2 (#82)

Writing & Preference benchmarks
BenchmarkGranite 4.2 30bQwen3.6 Plus
LMArena Text13611437
LMArena Creative Writing12881404
LMArena Multi-Turn13391438

Frequently asked questions

Is Granite 4.2 30b better than Qwen3.6 Plus?

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 41.8 on the Noometry Index.

Is Granite 4.2 30b or Qwen3.6 Plus better for coding?

They score almost the same on coding (41.0 vs 40.8); test both on your own repository before choosing.

How many benchmarks do Granite 4.2 30b and Qwen3.6 Plus share?

11 benchmarks have published results for both models. Granite 4.2 30b has 11 scored results on Noometry and Qwen3.6 Plus has 37.

Related comparisons

Go deeper