Model comparison

Llama 3.2 3B vs Qwen1.5 4b Chat

Llama 3.2 3B and Qwen1.5 4b Chat score almost the same on the Noometry Index (28.9 vs 28.8), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Llama 3.2 3B scores higher in 7 categories and Qwen1.5 4b Chat in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama 3.2 3B leads 56.0 to 49.0.

Side by side

Llama 3.2 3B and Qwen1.5 4b Chat specifications
Llama 3.2 3BQwen1.5 4b Chat
ProviderMetaAlibaba (Qwen)
Noometry Index28.928.8
Released2024-09-24—
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.33—
Results tracked1813

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5 4b Chat leads

Llama 3.2 3B: 27.6 (#319), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkLlama 3.2 3BQwen1.5 4b Chat
LMArena Coding1098999
BigCodeBench Instruct23.4%—
BigCodeBench Complete28.3%—

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Qwen1.5 4b Chat: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BQwen1.5 4b Chat
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Llama 3.2 3B leads

Llama 3.2 3B: 21.0 (#228), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkLlama 3.2 3BQwen1.5 4b Chat
LMArena Hard Prompts1095976

Math Llama 3.2 3B leads

Llama 3.2 3B: 32.4 (#214), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkLlama 3.2 3BQwen1.5 4b Chat
LMArena Math11261026

Knowledge Llama 3.2 3B leads

Llama 3.2 3B: 29.7 (#235), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkLlama 3.2 3BQwen1.5 4b Chat
LMArena Expert1090980

Multilingual Llama 3.2 3B leads

Llama 3.2 3B: 26.2 (#281), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkLlama 3.2 3BQwen1.5 4b Chat
LMArena Non-English1019979
LMArena Chinese10171024
LMArena German1056902
LMArena Russian949952

Instruction Following Llama 3.2 3B leads

Llama 3.2 3B: 56.0 (#275), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BQwen1.5 4b Chat
LMArena Instruction Following1089978

Long Context Llama 3.2 3B leads

Llama 3.2 3B: 33.4 (#261), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkLlama 3.2 3BQwen1.5 4b Chat
LMArena Longer Query1100988

Writing & Preference Too close to call

Llama 3.2 3B: 24.7 (#307), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BQwen1.5 4b Chat
LMArena Text1110997
LMArena Creative Writing1094969
LMArena Multi-Turn1105977
EQ-Bench Creative Writing595—

Frequently asked questions

Is Llama 3.2 3B better than Qwen1.5 4b Chat?

Llama 3.2 3B and Qwen1.5 4b Chat score almost the same on the Noometry Index (28.9 vs 28.8), so choose on price, context window or the category you care about most.

Is Llama 3.2 3B or Qwen1.5 4b Chat better for coding?

Qwen1.5 4b Chat scores higher on coding benchmarks: 29.1 versus 27.6 in the Noometry coding category.

How many benchmarks do Llama 3.2 3B and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper