Model comparison

Granite 3.0 2b Instruct vs Qwen3 8B

Qwen3 8B is the stronger model overall, scoring 33.7 to 30.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Granite 3.0 2b Instruct IBM

30.8

Rank #286 Confirmed

Qwen3 8B Alibaba (Qwen)

33.7

Rank #238 Confirmed

Summary

  • The widest gap is in knowledge, where Qwen3 8B leads 36.1 to 29.0.

Side by side

Granite 3.0 2b Instruct and Qwen3 8B specifications
Granite 3.0 2b InstructQwen3 8B
ProviderIBMAlibaba (Qwen)
Noometry Index30.833.7
Released—2025-04
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.18
Output $ / M tokens—$0.70
Results tracked1311

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 8B leads

Granite 3.0 2b Instruct: 28.3 (#316), Qwen3 8B: 34.0 (#248)

Coding benchmarks
BenchmarkGranite 3.0 2b InstructQwen3 8B
SciCode—22.6%
BigCodeBench Instruct20.5%—
LMArena Coding1090—

Agentic & Tool Use Not comparable

Granite 3.0 2b Instruct: —, Qwen3 8B: 30.2 (#78)

Agentic & Tool Use benchmarks
BenchmarkGranite 3.0 2b InstructQwen3 8B
Berkeley Function Calling Leaderboard—42.6%

Reasoning Granite 3.0 2b Instruct leads

Granite 3.0 2b Instruct: 20.5 (#235), Qwen3 8B: 16.6 (#303)

Reasoning benchmarks
BenchmarkGranite 3.0 2b InstructQwen3 8B
CritPt—0%
Chess Puzzles—5%
LMArena Hard Prompts1073—
DTBench—59.7%
LMCA—8.8%
Epoch Capabilities Index—136.17

Math Qwen3 8B leads

Granite 3.0 2b Instruct: 32.2 (#217), Qwen3 8B: 34.9 (#191)

Math benchmarks
BenchmarkGranite 3.0 2b InstructQwen3 8B
OTIS Mock AIME 2024-2025—56.1%
LMArena Math1117—

Knowledge Qwen3 8B leads

Granite 3.0 2b Instruct: 29.0 (#241), Qwen3 8B: 36.1 (#173)

Knowledge benchmarks
BenchmarkGranite 3.0 2b InstructQwen3 8B
GPQA Diamond—56.8%
Vectara Hallucination Rate—4.8%
LMArena Expert1064—

Multilingual Not comparable

Granite 3.0 2b Instruct: 27.0 (#278), Qwen3 8B: —

Multilingual benchmarks
BenchmarkGranite 3.0 2b InstructQwen3 8B
LMArena Non-English1033—
LMArena Chinese1070—
LMArena Russian1045—

Instruction Following Not comparable

Granite 3.0 2b Instruct: 53.9 (#284), Qwen3 8B: —

Instruction Following benchmarks
BenchmarkGranite 3.0 2b InstructQwen3 8B
LMArena Instruction Following1056—

Long Context Qwen3 8B leads

Granite 3.0 2b Instruct: 32.5 (#268), Qwen3 8B: 37.9 (#210)

Long Context benchmarks
BenchmarkGranite 3.0 2b InstructQwen3 8B
Fiction.LiveBench—62.1%
LMArena Longer Query1070—

Writing & Preference Not comparable

Granite 3.0 2b Instruct: 29.6 (#292), Qwen3 8B: —

Writing & Preference benchmarks
BenchmarkGranite 3.0 2b InstructQwen3 8B
LMArena Text1080—
LMArena Creative Writing1046—
LMArena Multi-Turn1053—

Frequently asked questions

Is Granite 3.0 2b Instruct better than Qwen3 8B?

Qwen3 8B is the stronger model overall, scoring 33.7 to 30.8 on the Noometry Index.

Is Granite 3.0 2b Instruct or Qwen3 8B better for coding?

Qwen3 8B scores higher on coding benchmarks: 34.0 versus 28.3 in the Noometry coding category.

How many benchmarks do Granite 3.0 2b Instruct and Qwen3 8B share?

0 benchmarks have published results for both models. Granite 3.0 2b Instruct has 13 scored results on Noometry and Qwen3 8B has 11.

Related comparisons

Go deeper