Model comparison

Granite 4.2 30b vs Llama 13b

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 24.4 on the Noometry Index.

Last verified . 7 shared benchmarks.

Granite 4.2 30b IBM

41.8

Rank #130 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Granite 4.2 30b scores higher in 5 categories and Llama 13b in 0 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Granite 4.2 30b leads 53.8 to 13.8.

Side by side

Granite 4.2 30b and Llama 13b specifications
Granite 4.2 30bLlama 13b
ProviderIBMMeta
Noometry Index41.824.4
Released—2023-02-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1121

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.2 30b leads

Granite 4.2 30b: 41.0 (#126), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkGranite 4.2 30bLlama 13b
LMArena Coding1396683

Reasoning Granite 4.2 30b leads

Granite 4.2 30b: 27.8 (#112), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkGranite 4.2 30bLlama 13b
LMArena Hard Prompts1374728
BIG-Bench Hard—37.9%
Epoch Capabilities Index—100.58
HellaSwag—79.2%
LAMBADA—75.2%
PIQA—80.1%
WinoGrande—73%

Math Not comparable

Granite 4.2 30b: —, Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkGranite 4.2 30bLlama 13b
LMArena Math—838
GSM8K—20.6%

Knowledge Not comparable

Granite 4.2 30b: 39.1 (#138), Llama 13b: —

Knowledge benchmarks
BenchmarkGranite 4.2 30bLlama 13b
LMArena Expert1406—
ARC (AI2) Challenge—52.7%
BoolQ—78.7%
MMLU—47.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

Granite 4.2 30b: —, Llama 13b: —

Multimodal benchmarks
BenchmarkGranite 4.2 30bLlama 13b
ScienceQA—43.3%

Multilingual Granite 4.2 30b leads

Granite 4.2 30b: 47.3 (#151), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkGranite 4.2 30bLlama 13b
LMArena Non-English1340819
LMArena Chinese1414—
LMArena Russian1343—

Instruction Following Granite 4.2 30b leads

Granite 4.2 30b: 71.2 (#155), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkGranite 4.2 30bLlama 13b
LMArena Instruction Following1347781

Long Context Not comparable

Granite 4.2 30b: 41.4 (#140), Llama 13b: —

Long Context benchmarks
BenchmarkGranite 4.2 30bLlama 13b
LMArena Longer Query1359—

Writing & Preference Granite 4.2 30b leads

Granite 4.2 30b: 53.8 (#156), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkGranite 4.2 30bLlama 13b
LMArena Text1361834
LMArena Creative Writing1288794
LMArena Multi-Turn1339753

Frequently asked questions

Is Granite 4.2 30b better than Llama 13b?

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 24.4 on the Noometry Index.

Is Granite 4.2 30b or Llama 13b better for coding?

Granite 4.2 30b scores higher on coding benchmarks: 41.0 versus 21.4 in the Noometry coding category.

How many benchmarks do Granite 4.2 30b and Llama 13b share?

7 benchmarks have published results for both models. Granite 4.2 30b has 11 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper