Model comparison

Llama 3.2 3B vs Qwen2.5-Coder (1.5B)

Llama 3.2 3B has enough public results to be ranked (#321); Qwen2.5-Coder (1.5B) does not yet, so treat this comparison as directional.

Last verified . 0 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Side by side

Llama 3.2 3B and Qwen2.5-Coder (1.5B) specifications
Llama 3.2 3BQwen2.5-Coder (1.5B)
ProviderMetaAlibaba (Qwen)
Noometry Index28.9—
Released2024-09-242024-09-18
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.33—
Results tracked186

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 3B: 27.6 (#319), Qwen2.5-Coder (1.5B): —

Coding benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder (1.5B)
BigCodeBench Instruct23.4%—
LMArena Coding1098—
BigCodeBench Complete28.3%—

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Qwen2.5-Coder (1.5B): —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder (1.5B)
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Not comparable

Llama 3.2 3B: 21.0 (#228), Qwen2.5-Coder (1.5B): —

Reasoning benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder (1.5B)
LMArena Hard Prompts1095—
Epoch Capabilities Index—113.14
HellaSwag—76.8%
WinoGrande—72.9%

Math Not comparable

Llama 3.2 3B: 32.4 (#214), Qwen2.5-Coder (1.5B): —

Math benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder (1.5B)
LMArena Math1126—
GSM8K—86.7%

Knowledge Not comparable

Llama 3.2 3B: 29.7 (#235), Qwen2.5-Coder (1.5B): —

Knowledge benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder (1.5B)
LMArena Expert1090—
ARC (AI2) Challenge—60.9%
MMLU—68%

Multilingual Not comparable

Llama 3.2 3B: 26.2 (#281), Qwen2.5-Coder (1.5B): —

Multilingual benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder (1.5B)
LMArena Non-English1019—
LMArena Chinese1017—
LMArena German1056—
LMArena Russian949—

Instruction Following Not comparable

Llama 3.2 3B: 56.0 (#275), Qwen2.5-Coder (1.5B): —

Instruction Following benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder (1.5B)
LMArena Instruction Following1089—

Long Context Not comparable

Llama 3.2 3B: 33.4 (#261), Qwen2.5-Coder (1.5B): —

Long Context benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder (1.5B)
LMArena Longer Query1100—

Writing & Preference Not comparable

Llama 3.2 3B: 24.7 (#307), Qwen2.5-Coder (1.5B): —

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BQwen2.5-Coder (1.5B)
LMArena Text1110—
LMArena Creative Writing1094—
EQ-Bench Creative Writing595—
LMArena Multi-Turn1105—

Frequently asked questions

Is Llama 3.2 3B better than Qwen2.5-Coder (1.5B)?

Llama 3.2 3B has enough public results to be ranked (#321); Qwen2.5-Coder (1.5B) does not yet, so treat this comparison as directional.

How many benchmarks do Llama 3.2 3B and Qwen2.5-Coder (1.5B) share?

0 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Qwen2.5-Coder (1.5B) has 6.

Related comparisons

Go deeper