Model comparison

Llama 3.2 1B vs Qwen2.5-Coder-32B

Qwen2.5-Coder-32B is the stronger model overall, scoring 33.4 to 20.1 on the Noometry Index. Llama 3.2 1B costs 11× less per token, which makes it the better buy when Qwen2.5-Coder-32B's lead doesn't matter for your workload.

Last verified . 15 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Llama 3.2 1B scores higher in 0 categories and Qwen2.5-Coder-32B in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen2.5-Coder-32B leads 33.4 to 7.2.
  • The biggest single-benchmark swing is BigCodeBench Complete: 11.3% for Llama 3.2 1B and 58% for Qwen2.5-Coder-32B.
  • Llama 3.2 1B is cheaper at $0.027 / $0.20 per million input/output tokens, against $0.66 / $1 for Qwen2.5-Coder-32B.
  • Llama 3.2 1B accepts more context: 60K tokens versus 33K.

Side by side

Llama 3.2 1B and Qwen2.5-Coder-32B specifications
Llama 3.2 1BQwen2.5-Coder-32B
ProviderMetaAlibaba (Qwen)
Noometry Index20.133.4
Released2024-09-242024-09-18
WeightsOpenOpen
Context window60K33K
Max output54K29K
Input $ / M tokens$0.027$0.66
Output $ / M tokens$0.20$1
Results tracked2231

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5-Coder-32B leads

Llama 3.2 1B: 21.1 (#338), Qwen2.5-Coder-32B: 22.6 (#333)

Coding benchmarks
BenchmarkLlama 3.2 1BQwen2.5-Coder-32B
BigCodeBench Instruct8.2%49%
LMArena Coding10701276
BigCodeBench Complete11.3%58%
SWE-bench Verified (bash only)—9%
Aider Polyglot—16.4%
LiveBench Coding—56.9%
HumanEval+—87.2%
MBPP+—77%

Agentic & Tool Use Not comparable

Llama 3.2 1B: 14.6 (#150), Qwen2.5-Coder-32B: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BQwen2.5-Coder-32B
Berkeley Function Calling Leaderboard10.8%—
BALROG6.6%—

Reasoning Qwen2.5-Coder-32B leads

Llama 3.2 1B: 16.2 (#308), Qwen2.5-Coder-32B: 21.2 (#225)

Reasoning benchmarks
BenchmarkLlama 3.2 1BQwen2.5-Coder-32B
LMArena Hard Prompts10441251
Epoch Capabilities Index101.99119.49
Chess Puzzles0%—
LiveBench Reasoning—42.1%
LiveBench Data Analysis—49.9%
HellaSwag—83%
LiveBench—46.2%
WinoGrande—80.8%

Math Qwen2.5-Coder-32B leads

Llama 3.2 1B: 10.4 (#313), Qwen2.5-Coder-32B: 33.3 (#204)

Math benchmarks
BenchmarkLlama 3.2 1BQwen2.5-Coder-32B
LMArena Math10861251
OTIS Mock AIME 2024-20250.6%—
LiveBench Math—46.6%
GSM8K—93%

Knowledge Qwen2.5-Coder-32B leads

Llama 3.2 1B: 7.2 (#312), Qwen2.5-Coder-32B: 33.4 (#203)

Knowledge benchmarks
BenchmarkLlama 3.2 1BQwen2.5-Coder-32B
LMArena Expert10071221
GPQA Diamond23.9%—
ARC (AI2) Challenge—70.5%
MMLU—79.1%

Multilingual Qwen2.5-Coder-32B leads

Llama 3.2 1B: 23.8 (#292), Qwen2.5-Coder-32B: 37.8 (#235)

Multilingual benchmarks
BenchmarkLlama 3.2 1BQwen2.5-Coder-32B
LMArena Non-English9731205
LMArena Chinese9591222
LMArena Russian9411228
LMArena German1014—

Instruction Following Qwen2.5-Coder-32B leads

Llama 3.2 1B: 52.4 (#290), Qwen2.5-Coder-32B: 61.4 (#245)

Instruction Following benchmarks
BenchmarkLlama 3.2 1BQwen2.5-Coder-32B
LMArena Instruction Following10311223
LiveBench Instruction Following—58.7%

Long Context Qwen2.5-Coder-32B leads

Llama 3.2 1B: 31.9 (#274), Qwen2.5-Coder-32B: 38.0 (#208)

Long Context benchmarks
BenchmarkLlama 3.2 1BQwen2.5-Coder-32B
LMArena Longer Query10501251

Writing & Preference Qwen2.5-Coder-32B leads

Llama 3.2 1B: 21.3 (#310), Qwen2.5-Coder-32B: 41.6 (#240)

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BQwen2.5-Coder-32B
LMArena Text10551230
LMArena Creative Writing10331174
LMArena Multi-Turn10301222
EQ-Bench Creative Writing200—
LiveBench Language—23.3%

Frequently asked questions

Is Llama 3.2 1B better than Qwen2.5-Coder-32B?

Qwen2.5-Coder-32B is the stronger model overall, scoring 33.4 to 20.1 on the Noometry Index. Llama 3.2 1B costs 11× less per token, which makes it the better buy when Qwen2.5-Coder-32B's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 1B or Qwen2.5-Coder-32B?

Llama 3.2 1B is cheaper. It lists at $0.027 per million input tokens and $0.20 per million output tokens; Qwen2.5-Coder-32B lists at $0.66 and $1.

Is Llama 3.2 1B or Qwen2.5-Coder-32B better for coding?

Qwen2.5-Coder-32B scores higher on coding benchmarks: 22.6 versus 21.1 in the Noometry coding category.

Which has the bigger context window?

Llama 3.2 1B does, with 60K tokens against 33K.

How many benchmarks do Llama 3.2 1B and Qwen2.5-Coder-32B share?

15 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Qwen2.5-Coder-32B has 31.

Related comparisons

Go deeper