Model comparison

Granite 3.1 2b Instruct vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 33.2 on the Noometry Index.

Last verified . 12 shared benchmarks.

Granite 3.1 2b Instruct IBM

33.2

Rank #247 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Granite 3.1 2b Instruct scores higher in 3 categories and Qwen Max in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Max leads 47.8 to 34.1.
  • Granite 3.1 2b Instruct has downloadable open weights; the other is API-only.

Side by side

Granite 3.1 2b Instruct and Qwen Max specifications
Granite 3.1 2b InstructQwen Max
ProviderIBMAlibaba (Qwen)
Noometry Index33.234.7
Released—2024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked1223

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 3.1 2b Instruct leads

Granite 3.1 2b Instruct: 33.4 (#257), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkGranite 3.1 2b InstructQwen Max
LMArena Coding11491288
Aider Polyglot—21.8%

Reasoning Qwen Max leads

Granite 3.1 2b Instruct: 22.0 (#209), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkGranite 3.1 2b InstructQwen Max
LMArena Hard Prompts11381269

Math Granite 3.1 2b Instruct leads

Granite 3.1 2b Instruct: 33.1 (#206), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkGranite 3.1 2b InstructQwen Max
LMArena Math11591275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Too close to call

Granite 3.1 2b Instruct: 30.8 (#224), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkGranite 3.1 2b InstructQwen Max
LMArena Expert11311248
GPQA Diamond—56.1%

Multilingual Qwen Max leads

Granite 3.1 2b Instruct: 29.1 (#269), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkGranite 3.1 2b InstructQwen Max
LMArena Non-English10681263
LMArena Chinese11391254
LMArena Russian10631274
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Spanish—1290

Instruction Following Qwen Max leads

Granite 3.1 2b Instruct: 57.7 (#264), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkGranite 3.1 2b InstructQwen Max
LMArena Instruction Following11161262

Long Context Qwen Max leads

Granite 3.1 2b Instruct: 35.0 (#244), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkGranite 3.1 2b InstructQwen Max
LMArena Longer Query11551288
Fiction.LiveBench—66.7%

Writing & Preference Qwen Max leads

Granite 3.1 2b Instruct: 34.1 (#274), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkGranite 3.1 2b InstructQwen Max
LMArena Text11271282
LMArena Creative Writing11161248
LMArena Multi-Turn10991277

Frequently asked questions

Is Granite 3.1 2b Instruct better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 33.2 on the Noometry Index.

Is Granite 3.1 2b Instruct or Qwen Max better for coding?

Granite 3.1 2b Instruct scores higher on coding benchmarks: 33.4 versus 30.7 in the Noometry coding category.

How many benchmarks do Granite 3.1 2b Instruct and Qwen Max share?

12 benchmarks have published results for both models. Granite 3.1 2b Instruct has 12 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper