Model comparison

Llama 3.2 3B vs Qwen2.5-Coder-32B

Qwen2.5-Coder-32B is the stronger model overall, scoring 33.4 to 28.9 on the Noometry Index. Llama 3.2 3B costs 6.2× less per token, which makes it the better buy when Qwen2.5-Coder-32B's lead doesn't matter for your workload.

Last verified . 14 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Llama 3.2 3B scores higher in 1 category and Qwen2.5-Coder-32B in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5-Coder-32B leads 41.6 to 24.7.
  • The biggest single-benchmark swing is BigCodeBench Complete: 28.3% for Llama 3.2 3B and 58% for Qwen2.5-Coder-32B.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $0.66 / $1 for Qwen2.5-Coder-32B.
  • Llama 3.2 3B accepts more context: 131K tokens versus 33K.

Side by side

Llama 3.2 3B and Qwen2.5-Coder-32B specifications
Llama 3.2 3BQwen2.5-Coder-32B
ProviderMetaAlibaba (Qwen)
Noometry Index28.933.4
Released2024-09-242024-09-18
WeightsOpenOpen
Context window131K33K
Max output118K29K
Input $ / M tokens$0.05$0.66
Output $ / M tokens$0.33$1
Results tracked1831

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.2 3B leads

Llama 3.2 3B: 27.6 (#319), Qwen2.5-Coder-32B: 22.6 (#333)

Coding benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder-32B
BigCodeBench Instruct23.4%49%
LMArena Coding10981276
BigCodeBench Complete28.3%58%
SWE-bench Verified (bash only)—9%
Aider Polyglot—16.4%
LiveBench Coding—56.9%
HumanEval+—87.2%
MBPP+—77%

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Qwen2.5-Coder-32B: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder-32B
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Too close to call

Llama 3.2 3B: 21.0 (#228), Qwen2.5-Coder-32B: 21.2 (#225)

Reasoning benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder-32B
LMArena Hard Prompts10951251
LiveBench Reasoning—42.1%
LiveBench Data Analysis—49.9%
Epoch Capabilities Index—119.49
HellaSwag—83%
LiveBench—46.2%
WinoGrande—80.8%

Math Too close to call

Llama 3.2 3B: 32.4 (#214), Qwen2.5-Coder-32B: 33.3 (#204)

Math benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder-32B
LMArena Math11261251
LiveBench Math—46.6%
GSM8K—93%

Knowledge Qwen2.5-Coder-32B leads

Llama 3.2 3B: 29.7 (#235), Qwen2.5-Coder-32B: 33.4 (#203)

Knowledge benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder-32B
LMArena Expert10901221
ARC (AI2) Challenge—70.5%
MMLU—79.1%

Multilingual Qwen2.5-Coder-32B leads

Llama 3.2 3B: 26.2 (#281), Qwen2.5-Coder-32B: 37.8 (#235)

Multilingual benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder-32B
LMArena Non-English10191205
LMArena Chinese10171222
LMArena Russian9491228
LMArena German1056—

Instruction Following Qwen2.5-Coder-32B leads

Llama 3.2 3B: 56.0 (#275), Qwen2.5-Coder-32B: 61.4 (#245)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder-32B
LMArena Instruction Following10891223
LiveBench Instruction Following—58.7%

Long Context Qwen2.5-Coder-32B leads

Llama 3.2 3B: 33.4 (#261), Qwen2.5-Coder-32B: 38.0 (#208)

Long Context benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder-32B
LMArena Longer Query11001251

Writing & Preference Qwen2.5-Coder-32B leads

Llama 3.2 3B: 24.7 (#307), Qwen2.5-Coder-32B: 41.6 (#240)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder-32B
LMArena Text11101230
LMArena Creative Writing10941174
LMArena Multi-Turn11051222
EQ-Bench Creative Writing595—
LiveBench Language—23.3%

Frequently asked questions

Is Llama 3.2 3B better than Qwen2.5-Coder-32B?

Qwen2.5-Coder-32B is the stronger model overall, scoring 33.4 to 28.9 on the Noometry Index. Llama 3.2 3B costs 6.2× less per token, which makes it the better buy when Qwen2.5-Coder-32B's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 3B or Qwen2.5-Coder-32B?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Qwen2.5-Coder-32B lists at $0.66 and $1.

Is Llama 3.2 3B or Qwen2.5-Coder-32B better for coding?

Llama 3.2 3B scores higher on coding benchmarks: 27.6 versus 22.6 in the Noometry coding category.

Which has the bigger context window?

Llama 3.2 3B does, with 131K tokens against 33K.

How many benchmarks do Llama 3.2 3B and Qwen2.5-Coder-32B share?

14 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Qwen2.5-Coder-32B has 31.

Related comparisons

Go deeper