Model comparison

Llama 3.2 1B vs Qwen3-1.7B

Qwen3-1.7B is the stronger model overall, scoring 26.6 to 20.1 on the Noometry Index.

Last verified . 4 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Qwen3-1.7B Alibaba (Qwen)

26.6

Rank #336 Reported

Summary

  • They share 4 benchmarks with published results for both. Llama 3.2 1B scores higher in 0 categories and Qwen3-1.7B in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3-1.7B leads 19.6 to 7.2.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 10.8% for Llama 3.2 1B and 28.4% for Qwen3-1.7B.

Side by side

Llama 3.2 1B and Qwen3-1.7B specifications
Llama 3.2 1BQwen3-1.7B
ProviderMetaAlibaba (Qwen)
Noometry Index20.126.6
Released2024-09-242025-04-29
WeightsOpenOpen
Context window60K—
Max output54K—
Input $ / M tokens$0.027—
Output $ / M tokens$0.20—
Results tracked224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 1B: 21.1 (#338), Qwen3-1.7B: —

Coding benchmarks
BenchmarkLlama 3.2 1BQwen3-1.7B
BigCodeBench Instruct8.2%—
LMArena Coding1070—
BigCodeBench Complete11.3%—

Agentic & Tool Use Qwen3-1.7B leads

Llama 3.2 1B: 14.6 (#150), Qwen3-1.7B: 24.7 (#115)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BQwen3-1.7B
Berkeley Function Calling Leaderboard10.8%28.4%
BALROG6.6%—

Reasoning Qwen3-1.7B leads

Llama 3.2 1B: 16.2 (#308), Qwen3-1.7B: 19.2 (#267)

Reasoning benchmarks
BenchmarkLlama 3.2 1BQwen3-1.7B
Chess Puzzles0%0%
LMArena Hard Prompts1044—
Epoch Capabilities Index101.99—

Math Qwen3-1.7B leads

Llama 3.2 1B: 10.4 (#313), Qwen3-1.7B: 16.3 (#294)

Math benchmarks
BenchmarkLlama 3.2 1BQwen3-1.7B
OTIS Mock AIME 2024-20250.6%8.1%
LMArena Math1086—

Knowledge Qwen3-1.7B leads

Llama 3.2 1B: 7.2 (#312), Qwen3-1.7B: 19.6 (#278)

Knowledge benchmarks
BenchmarkLlama 3.2 1BQwen3-1.7B
GPQA Diamond23.9%38%
LMArena Expert1007—

Multilingual Not comparable

Llama 3.2 1B: 23.8 (#292), Qwen3-1.7B: —

Multilingual benchmarks
BenchmarkLlama 3.2 1BQwen3-1.7B
LMArena Non-English973—
LMArena Chinese959—
LMArena German1014—
LMArena Russian941—

Instruction Following Not comparable

Llama 3.2 1B: 52.4 (#290), Qwen3-1.7B: —

Instruction Following benchmarks
BenchmarkLlama 3.2 1BQwen3-1.7B
LMArena Instruction Following1031—

Long Context Not comparable

Llama 3.2 1B: 31.9 (#274), Qwen3-1.7B: —

Long Context benchmarks
BenchmarkLlama 3.2 1BQwen3-1.7B
LMArena Longer Query1050—

Writing & Preference Not comparable

Llama 3.2 1B: 21.3 (#310), Qwen3-1.7B: —

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BQwen3-1.7B
LMArena Text1055—
LMArena Creative Writing1033—
EQ-Bench Creative Writing200—
LMArena Multi-Turn1030—

Frequently asked questions

Is Llama 3.2 1B better than Qwen3-1.7B?

Qwen3-1.7B is the stronger model overall, scoring 26.6 to 20.1 on the Noometry Index.

How many benchmarks do Llama 3.2 1B and Qwen3-1.7B share?

4 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Qwen3-1.7B has 4.

Related comparisons

Go deeper