Model comparison

Granite 4.2 30b vs Qwen2.5 7B Instruct

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 29.0 on the Noometry Index.

Last verified . 0 shared benchmarks.

Granite 4.2 30b IBM

41.8

Rank #130 Confirmed

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • The widest gap is in knowledge, where Granite 4.2 30b leads 39.1 to 17.0.

Side by side

Granite 4.2 30b and Qwen2.5 7B Instruct specifications
Granite 4.2 30bQwen2.5 7B Instruct
ProviderIBMAlibaba (Qwen)
Noometry Index41.829.0
Released—2024-09
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.17
Output $ / M tokens—$0.70
Results tracked1115

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.2 30b leads

Granite 4.2 30b: 41.0 (#126), Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
BenchmarkGranite 4.2 30bQwen2.5 7B Instruct
BigCodeBench Instruct—37.6%
LMArena Coding1396—
BigCodeBench Complete—46.1%

Agentic & Tool Use Not comparable

Granite 4.2 30b: —, Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
BenchmarkGranite 4.2 30bQwen2.5 7B Instruct
BALROG—7.8%

Reasoning Granite 4.2 30b leads

Granite 4.2 30b: 27.8 (#112), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
BenchmarkGranite 4.2 30bQwen2.5 7B Instruct
Chess Puzzles—0%
LMArena Hard Prompts1374—
DTBench—47.7%
LMCA—6.4%
Epoch Capabilities Index—118.51

Math Not comparable

Granite 4.2 30b: —, Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
BenchmarkGranite 4.2 30bQwen2.5 7B Instruct
OTIS Mock AIME 2024-2025—2.5%
Omni-MATH—29.4%

Knowledge Granite 4.2 30b leads

Granite 4.2 30b: 39.1 (#138), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
BenchmarkGranite 4.2 30bQwen2.5 7B Instruct
GPQA Diamond—35.5%
MMLU-Pro—53.9%
GPQA (HELM)—34.1%
LMArena Expert1406—
MMLU—72.9%

Multilingual Not comparable

Granite 4.2 30b: 47.3 (#151), Qwen2.5 7B Instruct: —

Multilingual benchmarks
BenchmarkGranite 4.2 30bQwen2.5 7B Instruct
LMArena Non-English1340—
LMArena Chinese1414—
LMArena Russian1343—

Instruction Following Granite 4.2 30b leads

Granite 4.2 30b: 71.2 (#155), Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
BenchmarkGranite 4.2 30bQwen2.5 7B Instruct
IFEval—74.1%
LMArena Instruction Following1347—

Long Context Not comparable

Granite 4.2 30b: 41.4 (#140), Qwen2.5 7B Instruct: —

Long Context benchmarks
BenchmarkGranite 4.2 30bQwen2.5 7B Instruct
LMArena Longer Query1359—

Writing & Preference Granite 4.2 30b leads

Granite 4.2 30b: 53.8 (#156), Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
BenchmarkGranite 4.2 30bQwen2.5 7B Instruct
LMArena Text1361—
LMArena Creative Writing1288—
WildBench—73.1%
LMArena Multi-Turn1339—

Frequently asked questions

Is Granite 4.2 30b better than Qwen2.5 7B Instruct?

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 29.0 on the Noometry Index.

Is Granite 4.2 30b or Qwen2.5 7B Instruct better for coding?

Granite 4.2 30b scores higher on coding benchmarks: 41.0 versus 36.5 in the Noometry coding category.

How many benchmarks do Granite 4.2 30b and Qwen2.5 7B Instruct share?

0 benchmarks have published results for both models. Granite 4.2 30b has 11 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper