Model comparison

Granite 3.1 8b Instruct vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 32.4 on the Noometry Index.

Last verified . 12 shared benchmarks.

Granite 3.1 8b Instruct IBM

32.4

Rank #258 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Granite 3.1 8b Instruct scores higher in 3 categories and Qwen Max in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Max leads 47.8 to 35.5.
  • Granite 3.1 8b Instruct has downloadable open weights; the other is API-only.

Side by side

Granite 3.1 8b Instruct and Qwen Max specifications
Granite 3.1 8b InstructQwen Max
ProviderIBMAlibaba (Qwen)
Noometry Index32.434.7
Released—2024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked1323

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 3.1 8b Instruct leads

Granite 3.1 8b Instruct: 34.5 (#233), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkGranite 3.1 8b InstructQwen Max
LMArena Coding11861288
Aider Polyglot—21.8%

Agentic & Tool Use Not comparable

Granite 3.1 8b Instruct: 24.1 (#120), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkGranite 3.1 8b InstructQwen Max
Berkeley Function Calling Leaderboard27.1%—

Reasoning Qwen Max leads

Granite 3.1 8b Instruct: 22.1 (#207), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkGranite 3.1 8b InstructQwen Max
LMArena Hard Prompts11451269

Math Granite 3.1 8b Instruct leads

Granite 3.1 8b Instruct: 33.0 (#209), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkGranite 3.1 8b InstructQwen Max
LMArena Math11521275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Too close to call

Granite 3.1 8b Instruct: 31.1 (#220), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkGranite 3.1 8b InstructQwen Max
LMArena Expert11421248
GPQA Diamond—56.1%

Multilingual Qwen Max leads

Granite 3.1 8b Instruct: 30.9 (#260), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkGranite 3.1 8b InstructQwen Max
LMArena Non-English10991263
LMArena Chinese11451254
LMArena Russian10921274
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Spanish—1290

Instruction Following Qwen Max leads

Granite 3.1 8b Instruct: 58.6 (#259), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkGranite 3.1 8b InstructQwen Max
LMArena Instruction Following11311262

Long Context Qwen Max leads

Granite 3.1 8b Instruct: 35.2 (#241), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkGranite 3.1 8b InstructQwen Max
LMArena Longer Query11621288
Fiction.LiveBench—66.7%

Writing & Preference Qwen Max leads

Granite 3.1 8b Instruct: 35.5 (#266), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkGranite 3.1 8b InstructQwen Max
LMArena Text11501282
LMArena Creative Writing11291248
LMArena Multi-Turn11081277

Frequently asked questions

Is Granite 3.1 8b Instruct better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 32.4 on the Noometry Index.

Is Granite 3.1 8b Instruct or Qwen Max better for coding?

Granite 3.1 8b Instruct scores higher on coding benchmarks: 34.5 versus 30.7 in the Noometry coding category.

How many benchmarks do Granite 3.1 8b Instruct and Qwen Max share?

12 benchmarks have published results for both models. Granite 3.1 8b Instruct has 13 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper