Model comparison

Llama 3.2 3B vs Olmo 3 32b Think

Olmo 3 32b Think is the stronger model overall, scoring 38.7 to 28.9 on the Noometry Index.

Last verified . 13 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Llama 3.2 3B scores higher in 0 categories and Olmo 3 32b Think in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Olmo 3 32b Think leads 49.1 to 24.7.

Side by side

Llama 3.2 3B and Olmo 3 32b Think specifications
Llama 3.2 3BOlmo 3 32b Think
ProviderMetaAllen Institute for AI (Ai2)
Noometry Index28.938.7
Released2024-09-24—
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.33—
Results tracked1814

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3 32b Think leads

Llama 3.2 3B: 27.6 (#319), Olmo 3 32b Think: 38.6 (#172)

Coding benchmarks
BenchmarkLlama 3.2 3BOlmo 3 32b Think
LMArena Coding10981319
BigCodeBench Instruct23.4%—
BigCodeBench Complete28.3%—

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Olmo 3 32b Think: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BOlmo 3 32b Think
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Olmo 3 32b Think leads

Llama 3.2 3B: 21.0 (#228), Olmo 3 32b Think: 25.9 (#140)

Reasoning benchmarks
BenchmarkLlama 3.2 3BOlmo 3 32b Think
LMArena Hard Prompts10951302

Math Olmo 3 32b Think leads

Llama 3.2 3B: 32.4 (#214), Olmo 3 32b Think: 36.5 (#165)

Math benchmarks
BenchmarkLlama 3.2 3BOlmo 3 32b Think
LMArena Math11261316

Knowledge Olmo 3 32b Think leads

Llama 3.2 3B: 29.7 (#235), Olmo 3 32b Think: 35.0 (#190)

Knowledge benchmarks
BenchmarkLlama 3.2 3BOlmo 3 32b Think
LMArena Expert10901273

Multilingual Olmo 3 32b Think leads

Llama 3.2 3B: 26.2 (#281), Olmo 3 32b Think: 41.2 (#210)

Multilingual benchmarks
BenchmarkLlama 3.2 3BOlmo 3 32b Think
LMArena Non-English10191255
LMArena Chinese10171300
LMArena German10561290
LMArena Russian9491254
LMArena French—1291

Instruction Following Olmo 3 32b Think leads

Llama 3.2 3B: 56.0 (#275), Olmo 3 32b Think: 67.2 (#198)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BOlmo 3 32b Think
LMArena Instruction Following10891275

Long Context Olmo 3 32b Think leads

Llama 3.2 3B: 33.4 (#261), Olmo 3 32b Think: 39.4 (#182)

Long Context benchmarks
BenchmarkLlama 3.2 3BOlmo 3 32b Think
LMArena Longer Query11001296

Writing & Preference Olmo 3 32b Think leads

Llama 3.2 3B: 24.7 (#307), Olmo 3 32b Think: 49.1 (#193)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BOlmo 3 32b Think
LMArena Text11101300
LMArena Creative Writing10941256
LMArena Multi-Turn11051290
EQ-Bench Creative Writing595—

Frequently asked questions

Is Llama 3.2 3B better than Olmo 3 32b Think?

Olmo 3 32b Think is the stronger model overall, scoring 38.7 to 28.9 on the Noometry Index.

Is Llama 3.2 3B or Olmo 3 32b Think better for coding?

Olmo 3 32b Think scores higher on coding benchmarks: 38.6 versus 27.6 in the Noometry coding category.

How many benchmarks do Llama 3.2 3B and Olmo 3 32b Think share?

13 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Olmo 3 32b Think has 14.

Related comparisons

Go deeper