Model comparison

Llama 3.2 3B vs Llama 4 Maverick

Llama 4 Maverick is the stronger model overall, scoring 30.9 to 28.9 on the Noometry Index. Llama 3.2 3B costs 2.5× less per token, which makes it the better buy when Llama 4 Maverick's lead doesn't matter for your workload.

Last verified . 17 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Llama 3.2 3B scores higher in 4 categories and Llama 4 Maverick in 5 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Llama 4 Maverick leads 42.2 to 26.2.
  • The biggest single-benchmark swing is BigCodeBench Complete: 28.3% for Llama 3.2 3B and 61.4% for Llama 4 Maverick.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $0.19 / $0.65 for Llama 4 Maverick.
  • Llama 3.2 3B accepts more context: 131K tokens versus 128K.

Side by side

Llama 3.2 3B and Llama 4 Maverick specifications
Llama 3.2 3BLlama 4 Maverick
ProviderMetaMeta
Noometry Index28.930.9
Released2024-09-242025-04-05
WeightsOpenOpen
Context window131K128K
Max output118K4K
Input $ / M tokens$0.05$0.19
Output $ / M tokens$0.33$0.65
Results tracked1854

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.2 3B leads

Llama 3.2 3B: 27.6 (#319), Llama 4 Maverick: 26.6 (#324)

Coding benchmarks
BenchmarkLlama 3.2 3BLlama 4 Maverick
BigCodeBench Instruct23.4%49.7%
LMArena Coding10981302
BigCodeBench Complete28.3%61.4%
SWE-bench Verified (bash only)—21%
Aider Polyglot—15.6%
SciCode—33.1%
WeirdML—24.5%
ALE-Bench—172.97

Agentic & Tool Use Llama 4 Maverick leads

Llama 3.2 3B: 20.1 (#143), Llama 4 Maverick: 28.2 (#91)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BLlama 4 Maverick
Berkeley Function Calling Leaderboard21.9%37.3%
BALROG10.1%—

Reasoning Llama 3.2 3B leads

Llama 3.2 3B: 21.0 (#228), Llama 4 Maverick: 10.1 (#342)

Reasoning benchmarks
BenchmarkLlama 3.2 3BLlama 4 Maverick
LMArena Hard Prompts10951281
ARC-AGI-2—0%
SimpleBench—27.7%
Kagi LLM Benchmark—55.9%
NYT Connections (extended)—8%
ARC-AGI-1—4.4%
CritPt—0%
EnigmaEval—0.6%
DTBench—61.9%
LMCA—15.9%
Epoch Capabilities Index—132.2
ForecastBench—57.5

Math Llama 3.2 3B leads

Llama 3.2 3B: 32.4 (#214), Llama 4 Maverick: 26.0 (#262)

Math benchmarks
BenchmarkLlama 3.2 3BLlama 4 Maverick
LMArena Math11261299
OTIS Mock AIME 2024-2025—20.6%
Omni-MATH—42.2%
MATH Level 5—73%
FrontierMath (Feb 2025 set)—0.7%

Knowledge Llama 4 Maverick leads

Llama 3.2 3B: 29.7 (#235), Llama 4 Maverick: 33.4 (#204)

Knowledge benchmarks
BenchmarkLlama 3.2 3BLlama 4 Maverick
LMArena Expert10901259
GPQA Diamond—67%
Humanity's Last Exam—5.7%
MMLU-Pro—81%
Confabulations—22.6%
Vectara Hallucination Rate—8.2%
GPQA (HELM)—65%

Multimodal Not comparable

Llama 3.2 3B: —, Llama 4 Maverick: 31.6 (#105)

Multimodal benchmarks
BenchmarkLlama 3.2 3BLlama 4 Maverick
LMArena Vision—1142
GeoBench—52%
SpatialViz-Bench—31.8%

Multilingual Llama 4 Maverick leads

Llama 3.2 3B: 26.2 (#281), Llama 4 Maverick: 42.2 (#195)

Multilingual benchmarks
BenchmarkLlama 3.2 3BLlama 4 Maverick
LMArena Non-English10191269
LMArena Chinese10171277
LMArena German10561291
LMArena Russian9491286
LMArena French—1259
LMArena Japanese—1207
LMArena Korean—1203
LMArena Spanish—1293

Instruction Following Llama 4 Maverick leads

Llama 3.2 3B: 56.0 (#275), Llama 4 Maverick: 71.7 (#146)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BLlama 4 Maverick
LMArena Instruction Following10891267
IFEval—90.8%

Long Context Llama 3.2 3B leads

Llama 3.2 3B: 33.4 (#261), Llama 4 Maverick: 31.4 (#279)

Long Context benchmarks
BenchmarkLlama 3.2 3BLlama 4 Maverick
LMArena Longer Query11001280
Fiction.LiveBench—46.2%

Writing & Preference Llama 4 Maverick leads

Llama 3.2 3B: 24.7 (#307), Llama 4 Maverick: 38.8 (#252)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BLlama 4 Maverick
LMArena Text11101287
LMArena Creative Writing10941267
EQ-Bench Creative Writing595860
LMArena Multi-Turn11051289
Short-Story Creative Writing—62%
WildBench—80%

Frequently asked questions

Is Llama 3.2 3B better than Llama 4 Maverick?

Llama 4 Maverick is the stronger model overall, scoring 30.9 to 28.9 on the Noometry Index. Llama 3.2 3B costs 2.5× less per token, which makes it the better buy when Llama 4 Maverick's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 3B or Llama 4 Maverick?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Llama 4 Maverick lists at $0.19 and $0.65.

Is Llama 3.2 3B or Llama 4 Maverick better for coding?

Llama 3.2 3B scores higher on coding benchmarks: 27.6 versus 26.6 in the Noometry coding category.

Which has the bigger context window?

Llama 3.2 3B does, with 131K tokens against 128K.

How many benchmarks do Llama 3.2 3B and Llama 4 Maverick share?

17 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Llama 4 Maverick has 54.

Related comparisons

Go deeper