Model comparison

Llama 3.2 3B vs Llama-3.3-70B-Instruct

Llama-3.3-70B-Instruct is the stronger model overall, scoring 30.6 to 28.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Llama 3.2 3B scores higher in 3 categories and Llama-3.3-70B-Instruct in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama-3.3-70B-Instruct leads 47.6 to 24.7.
  • The biggest single-benchmark swing is BigCodeBench Complete: 28.3% for Llama 3.2 3B and 57.5% for Llama-3.3-70B-Instruct.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $0.10 / $0.32 for Llama-3.3-70B-Instruct.
  • Llama 3.2 3B accepts more context: 131K tokens versus 128K.

Side by side

Llama 3.2 3B and Llama-3.3-70B-Instruct specifications
Llama 3.2 3BLlama-3.3-70B-Instruct
ProviderMetaMeta
Noometry Index28.930.6
Released2024-09-242024-12-06
WeightsOpenOpen
Context window131K128K
Max output118K4K
Input $ / M tokens$0.05$0.10
Output $ / M tokens$0.33$0.32
Results tracked1843

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama-3.3-70B-Instruct leads

Llama 3.2 3B: 27.6 (#319), Llama-3.3-70B-Instruct: 31.0 (#290)

Coding benchmarks
BenchmarkLlama 3.2 3BLlama-3.3-70B-Instruct
BigCodeBench Instruct23.4%46.9%
LMArena Coding10981268
BigCodeBench Complete28.3%57.5%
SciCode—26%
WeirdML—14.4%
LiveBench Coding—36.6%

Agentic & Tool Use Llama-3.3-70B-Instruct leads

Llama 3.2 3B: 20.1 (#143), Llama-3.3-70B-Instruct: 25.8 (#105)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BLlama-3.3-70B-Instruct
Berkeley Function Calling Leaderboard21.9%31.9%
BALROG10.1%23%

Reasoning Llama 3.2 3B leads

Llama 3.2 3B: 21.0 (#228), Llama-3.3-70B-Instruct: 14.1 (#327)

Reasoning benchmarks
BenchmarkLlama 3.2 3BLlama-3.3-70B-Instruct
LMArena Hard Prompts10951257
SimpleBench—19.9%
CritPt—0%
LiveBench Reasoning—50.8%
DTBench—59.5%
LiveBench Data Analysis—49.5%
LMCA—17.5%
Epoch Capabilities Index—127.33
ForecastBench—58.6
LiveBench—50.2%

Math Llama 3.2 3B leads

Llama 3.2 3B: 32.4 (#214), Llama-3.3-70B-Instruct: 15.3 (#298)

Math benchmarks
BenchmarkLlama 3.2 3BLlama-3.3-70B-Instruct
LMArena Math11261267
OTIS Mock AIME 2024-2025—5.1%
LiveBench Math—42.2%
MATH Level 5—41.6%

Knowledge Too close to call

Llama 3.2 3B: 29.7 (#235), Llama-3.3-70B-Instruct: 30.6 (#226)

Knowledge benchmarks
BenchmarkLlama 3.2 3BLlama-3.3-70B-Instruct
LMArena Expert10901225
GPQA Diamond—47.4%
Confabulations—22.8%
Vectara Hallucination Rate—4.1%
MMLU—86.3%

Multilingual Llama-3.3-70B-Instruct leads

Llama 3.2 3B: 26.2 (#281), Llama-3.3-70B-Instruct: 39.9 (#220)

Multilingual benchmarks
BenchmarkLlama 3.2 3BLlama-3.3-70B-Instruct
LMArena Non-English10191236
LMArena Chinese10171217
LMArena German10561251
LMArena Russian9491252
LMArena French—1281
LMArena Japanese—1150
LMArena Korean—1143
LMArena Spanish—1270

Instruction Following Llama-3.3-70B-Instruct leads

Llama 3.2 3B: 56.0 (#275), Llama-3.3-70B-Instruct: 71.1 (#157)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BLlama-3.3-70B-Instruct
LMArena Instruction Following10891242
LiveBench Instruction Following—82.7%

Long Context Llama 3.2 3B leads

Llama 3.2 3B: 33.4 (#261), Llama-3.3-70B-Instruct: 26.4 (#295)

Long Context benchmarks
BenchmarkLlama 3.2 3BLlama-3.3-70B-Instruct
LMArena Longer Query11001256
Fiction.LiveBench—33.3%

Writing & Preference Llama-3.3-70B-Instruct leads

Llama 3.2 3B: 24.7 (#307), Llama-3.3-70B-Instruct: 47.6 (#207)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BLlama-3.3-70B-Instruct
LMArena Text11101274
LMArena Creative Writing10941250
LMArena Multi-Turn11051280
EQ-Bench Creative Writing595—
LiveBench Language—39.2%

Frequently asked questions

Is Llama 3.2 3B better than Llama-3.3-70B-Instruct?

Llama-3.3-70B-Instruct is the stronger model overall, scoring 30.6 to 28.9 on the Noometry Index.

Which is cheaper, Llama 3.2 3B or Llama-3.3-70B-Instruct?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Llama-3.3-70B-Instruct lists at $0.10 and $0.32.

Is Llama 3.2 3B or Llama-3.3-70B-Instruct better for coding?

Llama-3.3-70B-Instruct scores higher on coding benchmarks: 31.0 versus 27.6 in the Noometry coding category.

Which has the bigger context window?

Llama 3.2 3B does, with 131K tokens against 128K.

How many benchmarks do Llama 3.2 3B and Llama-3.3-70B-Instruct share?

17 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Llama-3.3-70B-Instruct has 43.

Related comparisons

Go deeper