Model comparison

Llama 3.2 3B vs Qwen1.5-110B

Qwen1.5-110B is the stronger model overall, scoring 34.2 to 28.9 on the Noometry Index.

Last verified . 15 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Qwen1.5-110B Alibaba (Qwen)

34.2

Rank #234 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Llama 3.2 3B scores higher in 0 categories and Qwen1.5-110B in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen1.5-110B leads 38.0 to 24.7.
  • The biggest single-benchmark swing is BigCodeBench Complete: 28.3% for Llama 3.2 3B and 44.4% for Qwen1.5-110B.

Side by side

Llama 3.2 3B and Qwen1.5-110B specifications
Llama 3.2 3BQwen1.5-110B
ProviderMetaAlibaba (Qwen)
Noometry Index28.934.2
Released2024-09-242024-04-25
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.33—
Results tracked1820

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5-110B leads

Llama 3.2 3B: 27.6 (#319), Qwen1.5-110B: 33.0 (#264)

Coding benchmarks
BenchmarkLlama 3.2 3BQwen1.5-110B
BigCodeBench Instruct23.4%35%
LMArena Coding10981184
BigCodeBench Complete28.3%44.4%

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Qwen1.5-110B: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BQwen1.5-110B
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Qwen1.5-110B leads

Llama 3.2 3B: 21.0 (#228), Qwen1.5-110B: 22.7 (#189)

Reasoning benchmarks
BenchmarkLlama 3.2 3BQwen1.5-110B
LMArena Hard Prompts10951168
ForecastBench—57.7

Math Qwen1.5-110B leads

Llama 3.2 3B: 32.4 (#214), Qwen1.5-110B: 33.7 (#201)

Math benchmarks
BenchmarkLlama 3.2 3BQwen1.5-110B
LMArena Math11261185

Knowledge Qwen1.5-110B leads

Llama 3.2 3B: 29.7 (#235), Qwen1.5-110B: 31.2 (#219)

Knowledge benchmarks
BenchmarkLlama 3.2 3BQwen1.5-110B
LMArena Expert10901144

Multilingual Qwen1.5-110B leads

Llama 3.2 3B: 26.2 (#281), Qwen1.5-110B: 33.6 (#250)

Multilingual benchmarks
BenchmarkLlama 3.2 3BQwen1.5-110B
LMArena Non-English10191142
LMArena Chinese10171206
LMArena German10561123
LMArena Russian9491118
LMArena French—1151
LMArena Japanese—1074
LMArena Korean—1044
LMArena Spanish—1142

Instruction Following Qwen1.5-110B leads

Llama 3.2 3B: 56.0 (#275), Qwen1.5-110B: 60.3 (#252)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BQwen1.5-110B
LMArena Instruction Following10891158

Long Context Qwen1.5-110B leads

Llama 3.2 3B: 33.4 (#261), Qwen1.5-110B: 35.1 (#242)

Long Context benchmarks
BenchmarkLlama 3.2 3BQwen1.5-110B
LMArena Longer Query11001157

Writing & Preference Qwen1.5-110B leads

Llama 3.2 3B: 24.7 (#307), Qwen1.5-110B: 38.0 (#255)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BQwen1.5-110B
LMArena Text11101175
LMArena Creative Writing10941148
LMArena Multi-Turn11051160
EQ-Bench Creative Writing595—

Frequently asked questions

Is Llama 3.2 3B better than Qwen1.5-110B?

Qwen1.5-110B is the stronger model overall, scoring 34.2 to 28.9 on the Noometry Index.

Is Llama 3.2 3B or Qwen1.5-110B better for coding?

Qwen1.5-110B scores higher on coding benchmarks: 33.0 versus 27.6 in the Noometry coding category.

How many benchmarks do Llama 3.2 3B and Qwen1.5-110B share?

15 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Qwen1.5-110B has 20.

Related comparisons

Go deeper