Model comparison

Llama 3.2 3B vs Qwen Plus

Qwen Plus is the stronger model overall, scoring 37.1 to 28.9 on the Noometry Index. Llama 3.2 3B costs 5.0× less per token, which makes it the better buy when Qwen Plus's lead doesn't matter for your workload.

Last verified . 12 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Llama 3.2 3B scores higher in 2 categories and Qwen Plus in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Plus leads 52.2 to 24.7.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $0.40 / $1.20 for Qwen Plus.
  • Qwen Plus accepts more context: 1M tokens versus 131K.
  • Llama 3.2 3B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 3B and Qwen Plus specifications
Llama 3.2 3BQwen Plus
ProviderMetaAlibaba (Qwen)
Noometry Index28.937.1
Released2024-09-242024-01-25
WeightsOpenProprietary
Context window131K1M
Max output118K33K
Input $ / M tokens$0.05$0.40
Output $ / M tokens$0.33$1.20
Results tracked1820

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Plus leads

Llama 3.2 3B: 27.6 (#319), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkLlama 3.2 3BQwen Plus
LMArena Coding10981328
BigCodeBench Instruct23.4%—
BigCodeBench Complete28.3%—

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Qwen Plus: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BQwen Plus
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Qwen Plus leads

Llama 3.2 3B: 21.0 (#228), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkLlama 3.2 3BQwen Plus
LMArena Hard Prompts10951317
Kagi LLM Benchmark—63.3%
DTBench—81.1%
LMCA—24%

Math Llama 3.2 3B leads

Llama 3.2 3B: 32.4 (#214), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkLlama 3.2 3BQwen Plus
LMArena Math11261326
OTIS Mock AIME 2024-2025—17.8%
MATH Level 5—65.3%
FrontierMath (Feb 2025 set)—1.7%

Knowledge Llama 3.2 3B leads

Llama 3.2 3B: 29.7 (#235), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkLlama 3.2 3BQwen Plus
LMArena Expert10901328
GPQA Diamond—48.1%

Multilingual Qwen Plus leads

Llama 3.2 3B: 26.2 (#281), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkLlama 3.2 3BQwen Plus
LMArena Non-English10191310
LMArena Chinese10171347
LMArena Russian9491323
LMArena German1056—
LMArena Japanese—1251

Instruction Following Qwen Plus leads

Llama 3.2 3B: 56.0 (#275), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BQwen Plus
LMArena Instruction Following10891303

Long Context Qwen Plus leads

Llama 3.2 3B: 33.4 (#261), Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkLlama 3.2 3BQwen Plus
LMArena Longer Query11001324

Writing & Preference Qwen Plus leads

Llama 3.2 3B: 24.7 (#307), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BQwen Plus
LMArena Text11101326
LMArena Creative Writing10941293
LMArena Multi-Turn11051336
EQ-Bench Creative Writing595—

Frequently asked questions

Is Llama 3.2 3B better than Qwen Plus?

Qwen Plus is the stronger model overall, scoring 37.1 to 28.9 on the Noometry Index. Llama 3.2 3B costs 5.0× less per token, which makes it the better buy when Qwen Plus's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 3B or Qwen Plus?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Qwen Plus lists at $0.40 and $1.20.

Is Llama 3.2 3B or Qwen Plus better for coding?

Qwen Plus scores higher on coding benchmarks: 38.9 versus 27.6 in the Noometry coding category.

Which has the bigger context window?

Qwen Plus does, with 1M tokens against 131K.

How many benchmarks do Llama 3.2 3B and Qwen Plus share?

12 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper