Model comparison

Llama 3.2 3B vs Olmo 2 0325 32b Instruct

Olmo 2 0325 32b Instruct is the stronger model overall, scoring 32.7 to 28.9 on the Noometry Index.

Last verified . 11 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Llama 3.2 3B scores higher in 2 categories and Olmo 2 0325 32b Instruct in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Olmo 2 0325 32b Instruct leads 42.1 to 24.7.

Side by side

Llama 3.2 3B and Olmo 2 0325 32b Instruct specifications
Llama 3.2 3BOlmo 2 0325 32b Instruct
ProviderMetaAllen Institute for AI (Ai2)
Noometry Index28.932.7
Released2024-09-24—
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.33—
Results tracked1816

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 2 0325 32b Instruct leads

Llama 3.2 3B: 27.6 (#319), Olmo 2 0325 32b Instruct: 35.2 (#227)

Coding benchmarks
BenchmarkLlama 3.2 3BOlmo 2 0325 32b Instruct
LMArena Coding10981210
BigCodeBench Instruct23.4%—
BigCodeBench Complete28.3%—

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Olmo 2 0325 32b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BOlmo 2 0325 32b Instruct
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Olmo 2 0325 32b Instruct leads

Llama 3.2 3B: 21.0 (#228), Olmo 2 0325 32b Instruct: 23.6 (#175)

Reasoning benchmarks
BenchmarkLlama 3.2 3BOlmo 2 0325 32b Instruct
LMArena Hard Prompts10951208

Math Llama 3.2 3B leads

Llama 3.2 3B: 32.4 (#214), Olmo 2 0325 32b Instruct: 26.8 (#255)

Math benchmarks
BenchmarkLlama 3.2 3BOlmo 2 0325 32b Instruct
LMArena Math11261208
Omni-MATH—16.1%

Knowledge Llama 3.2 3B leads

Llama 3.2 3B: 29.7 (#235), Olmo 2 0325 32b Instruct: 19.5 (#279)

Knowledge benchmarks
BenchmarkLlama 3.2 3BOlmo 2 0325 32b Instruct
MMLU-Pro—41.4%
GPQA (HELM)—28.7%
LMArena Expert1090—

Multilingual Olmo 2 0325 32b Instruct leads

Llama 3.2 3B: 26.2 (#281), Olmo 2 0325 32b Instruct: 34.8 (#248)

Multilingual benchmarks
BenchmarkLlama 3.2 3BOlmo 2 0325 32b Instruct
LMArena Non-English10191160
LMArena Chinese10171192
LMArena Russian9491187
LMArena German1056—

Instruction Following Olmo 2 0325 32b Instruct leads

Llama 3.2 3B: 56.0 (#275), Olmo 2 0325 32b Instruct: 61.5 (#244)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BOlmo 2 0325 32b Instruct
LMArena Instruction Following10891186
IFEval—78%

Long Context Olmo 2 0325 32b Instruct leads

Llama 3.2 3B: 33.4 (#261), Olmo 2 0325 32b Instruct: 36.2 (#234)

Long Context benchmarks
BenchmarkLlama 3.2 3BOlmo 2 0325 32b Instruct
LMArena Longer Query11001194

Writing & Preference Olmo 2 0325 32b Instruct leads

Llama 3.2 3B: 24.7 (#307), Olmo 2 0325 32b Instruct: 42.1 (#236)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BOlmo 2 0325 32b Instruct
LMArena Text11101218
LMArena Creative Writing10941199
LMArena Multi-Turn11051221
EQ-Bench Creative Writing595—
WildBench—73.4%

Frequently asked questions

Is Llama 3.2 3B better than Olmo 2 0325 32b Instruct?

Olmo 2 0325 32b Instruct is the stronger model overall, scoring 32.7 to 28.9 on the Noometry Index.

Is Llama 3.2 3B or Olmo 2 0325 32b Instruct better for coding?

Olmo 2 0325 32b Instruct scores higher on coding benchmarks: 35.2 versus 27.6 in the Noometry coding category.

How many benchmarks do Llama 3.2 3B and Olmo 2 0325 32b Instruct share?

11 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Olmo 2 0325 32b Instruct has 16.

Related comparisons

Go deeper