Model comparison

Llama 3.2 3B vs Qwen2.5 7B Instruct

Llama 3.2 3B and Qwen2.5 7B Instruct score almost the same on the Noometry Index (28.9 vs 29.0), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Llama 3.2 3B scores higher in 3 categories and Qwen2.5 7B Instruct in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5 7B Instruct leads 48.8 to 24.7.
  • The biggest single-benchmark swing is BigCodeBench Complete: 28.3% for Llama 3.2 3B and 46.1% for Qwen2.5 7B Instruct.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $0.17 / $0.70 for Qwen2.5 7B Instruct.

Side by side

Llama 3.2 3B and Qwen2.5 7B Instruct specifications
Llama 3.2 3BQwen2.5 7B Instruct
ProviderMetaAlibaba (Qwen)
Noometry Index28.929.0
Released2024-09-242024-09
WeightsOpenOpen
Context window131K131K
Max output118K8K
Input $ / M tokens$0.05$0.17
Output $ / M tokens$0.33$0.70
Results tracked1815

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 7B Instruct leads

Llama 3.2 3B: 27.6 (#319), Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
BenchmarkLlama 3.2 3BQwen2.5 7B Instruct
BigCodeBench Instruct23.4%37.6%
BigCodeBench Complete28.3%46.1%
LMArena Coding1098—

Agentic & Tool Use Qwen2.5 7B Instruct leads

Llama 3.2 3B: 20.1 (#143), Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BQwen2.5 7B Instruct
BALROG10.1%7.8%
Berkeley Function Calling Leaderboard21.9%—

Reasoning Llama 3.2 3B leads

Llama 3.2 3B: 21.0 (#228), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
BenchmarkLlama 3.2 3BQwen2.5 7B Instruct
Chess Puzzles—0%
LMArena Hard Prompts1095—
DTBench—47.7%
LMCA—6.4%
Epoch Capabilities Index—118.51

Math Llama 3.2 3B leads

Llama 3.2 3B: 32.4 (#214), Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
BenchmarkLlama 3.2 3BQwen2.5 7B Instruct
OTIS Mock AIME 2024-2025—2.5%
Omni-MATH—29.4%
LMArena Math1126—

Knowledge Llama 3.2 3B leads

Llama 3.2 3B: 29.7 (#235), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
BenchmarkLlama 3.2 3BQwen2.5 7B Instruct
GPQA Diamond—35.5%
MMLU-Pro—53.9%
GPQA (HELM)—34.1%
LMArena Expert1090—
MMLU—72.9%

Multilingual Not comparable

Llama 3.2 3B: 26.2 (#281), Qwen2.5 7B Instruct: —

Multilingual benchmarks
BenchmarkLlama 3.2 3BQwen2.5 7B Instruct
LMArena Non-English1019—
LMArena Chinese1017—
LMArena German1056—
LMArena Russian949—

Instruction Following Qwen2.5 7B Instruct leads

Llama 3.2 3B: 56.0 (#275), Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BQwen2.5 7B Instruct
IFEval—74.1%
LMArena Instruction Following1089—

Long Context Not comparable

Llama 3.2 3B: 33.4 (#261), Qwen2.5 7B Instruct: —

Long Context benchmarks
BenchmarkLlama 3.2 3BQwen2.5 7B Instruct
LMArena Longer Query1100—

Writing & Preference Qwen2.5 7B Instruct leads

Llama 3.2 3B: 24.7 (#307), Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BQwen2.5 7B Instruct
LMArena Text1110—
LMArena Creative Writing1094—
EQ-Bench Creative Writing595—
WildBench—73.1%
LMArena Multi-Turn1105—

Frequently asked questions

Is Llama 3.2 3B better than Qwen2.5 7B Instruct?

Llama 3.2 3B and Qwen2.5 7B Instruct score almost the same on the Noometry Index (28.9 vs 29.0), so choose on price, context window or the category you care about most.

Which is cheaper, Llama 3.2 3B or Qwen2.5 7B Instruct?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Qwen2.5 7B Instruct lists at $0.17 and $0.70.

Is Llama 3.2 3B or Qwen2.5 7B Instruct better for coding?

Qwen2.5 7B Instruct scores higher on coding benchmarks: 36.5 versus 27.6 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Llama 3.2 3B and Qwen2.5 7B Instruct share?

3 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper