Model comparison

Llama 3.2 1B vs Qwen2.5 7B Instruct

Qwen2.5 7B Instruct is the stronger model overall, scoring 29.0 to 20.1 on the Noometry Index. Llama 3.2 1B costs 4.3× less per token, which makes it the better buy when Qwen2.5 7B Instruct's lead doesn't matter for your workload.

Last verified . 7 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Llama 3.2 1B scores higher in 1 category and Qwen2.5 7B Instruct in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5 7B Instruct leads 48.8 to 21.3.
  • The biggest single-benchmark swing is BigCodeBench Complete: 11.3% for Llama 3.2 1B and 46.1% for Qwen2.5 7B Instruct.
  • Llama 3.2 1B is cheaper at $0.027 / $0.20 per million input/output tokens, against $0.17 / $0.70 for Qwen2.5 7B Instruct.
  • Qwen2.5 7B Instruct accepts more context: 131K tokens versus 60K.

Side by side

Llama 3.2 1B and Qwen2.5 7B Instruct specifications
Llama 3.2 1BQwen2.5 7B Instruct
ProviderMetaAlibaba (Qwen)
Noometry Index20.129.0
Released2024-09-242024-09
WeightsOpenOpen
Context window60K131K
Max output54K8K
Input $ / M tokens$0.027$0.17
Output $ / M tokens$0.20$0.70
Results tracked2215

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 7B Instruct leads

Llama 3.2 1B: 21.1 (#338), Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
BenchmarkLlama 3.2 1BQwen2.5 7B Instruct
BigCodeBench Instruct8.2%37.6%
BigCodeBench Complete11.3%46.1%
LMArena Coding1070—

Agentic & Tool Use Qwen2.5 7B Instruct leads

Llama 3.2 1B: 14.6 (#150), Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BQwen2.5 7B Instruct
BALROG6.6%7.8%
Berkeley Function Calling Leaderboard10.8%—

Reasoning Llama 3.2 1B leads

Llama 3.2 1B: 16.2 (#308), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
BenchmarkLlama 3.2 1BQwen2.5 7B Instruct
Chess Puzzles0%0%
Epoch Capabilities Index101.99118.51
LMArena Hard Prompts1044—
DTBench—47.7%
LMCA—6.4%

Math Qwen2.5 7B Instruct leads

Llama 3.2 1B: 10.4 (#313), Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
BenchmarkLlama 3.2 1BQwen2.5 7B Instruct
OTIS Mock AIME 2024-20250.6%2.5%
Omni-MATH—29.4%
LMArena Math1086—

Knowledge Qwen2.5 7B Instruct leads

Llama 3.2 1B: 7.2 (#312), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
BenchmarkLlama 3.2 1BQwen2.5 7B Instruct
GPQA Diamond23.9%35.5%
MMLU-Pro—53.9%
GPQA (HELM)—34.1%
LMArena Expert1007—
MMLU—72.9%

Multilingual Not comparable

Llama 3.2 1B: 23.8 (#292), Qwen2.5 7B Instruct: —

Multilingual benchmarks
BenchmarkLlama 3.2 1BQwen2.5 7B Instruct
LMArena Non-English973—
LMArena Chinese959—
LMArena German1014—
LMArena Russian941—

Instruction Following Qwen2.5 7B Instruct leads

Llama 3.2 1B: 52.4 (#290), Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
BenchmarkLlama 3.2 1BQwen2.5 7B Instruct
IFEval—74.1%
LMArena Instruction Following1031—

Long Context Not comparable

Llama 3.2 1B: 31.9 (#274), Qwen2.5 7B Instruct: —

Long Context benchmarks
BenchmarkLlama 3.2 1BQwen2.5 7B Instruct
LMArena Longer Query1050—

Writing & Preference Qwen2.5 7B Instruct leads

Llama 3.2 1B: 21.3 (#310), Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BQwen2.5 7B Instruct
LMArena Text1055—
LMArena Creative Writing1033—
EQ-Bench Creative Writing200—
WildBench—73.1%
LMArena Multi-Turn1030—

Frequently asked questions

Is Llama 3.2 1B better than Qwen2.5 7B Instruct?

Qwen2.5 7B Instruct is the stronger model overall, scoring 29.0 to 20.1 on the Noometry Index. Llama 3.2 1B costs 4.3× less per token, which makes it the better buy when Qwen2.5 7B Instruct's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 1B or Qwen2.5 7B Instruct?

Llama 3.2 1B is cheaper. It lists at $0.027 per million input tokens and $0.20 per million output tokens; Qwen2.5 7B Instruct lists at $0.17 and $0.70.

Is Llama 3.2 1B or Qwen2.5 7B Instruct better for coding?

Qwen2.5 7B Instruct scores higher on coding benchmarks: 36.5 versus 21.1 in the Noometry coding category.

Which has the bigger context window?

Qwen2.5 7B Instruct does, with 131K tokens against 60K.

How many benchmarks do Llama 3.2 1B and Qwen2.5 7B Instruct share?

7 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper