Model comparison

Llama 3.2 90B vs Qwen2.5 7B Instruct

Qwen2.5 7B Instruct is the stronger model overall, scoring 29.0 to 27.5 on the Noometry Index.

Last verified . 5 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Llama 3.2 90B scores higher in 3 categories and Qwen2.5 7B Instruct in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Llama 3.2 90B leads 21.7 to 14.8.
  • The biggest single-benchmark swing is BALROG: 27.3% for Llama 3.2 90B and 7.8% for Qwen2.5 7B Instruct.

Side by side

Llama 3.2 90B and Qwen2.5 7B Instruct specifications
Llama 3.2 90BQwen2.5 7B Instruct
ProviderMetaAlibaba (Qwen)
Noometry Index27.529.0
Released2024-09-242024-09
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.17
Output $ / M tokens—$0.70
Results tracked915

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
BenchmarkLlama 3.2 90BQwen2.5 7B Instruct
BigCodeBench Instruct—37.6%
BigCodeBench Complete—46.1%

Agentic & Tool Use Llama 3.2 90B leads

Llama 3.2 90B: 30.0 (#80), Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BQwen2.5 7B Instruct
BALROG27.3%7.8%

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
BenchmarkLlama 3.2 90BQwen2.5 7B Instruct
Epoch Capabilities Index125.5118.51
Chess Puzzles—0%
EnigmaEval0.4%—
DTBench—47.7%
LMCA—6.4%

Math Qwen2.5 7B Instruct leads

Llama 3.2 90B: 11.1 (#308), Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
BenchmarkLlama 3.2 90BQwen2.5 7B Instruct
OTIS Mock AIME 2024-20252.6%2.5%
Omni-MATH—29.4%
MATH Level 539.4%—

Knowledge Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#274), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
BenchmarkLlama 3.2 90BQwen2.5 7B Instruct
GPQA Diamond41%35.5%
MMLU80.3%72.9%
MMLU-Pro—53.9%
GPQA (HELM)—34.1%

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Qwen2.5 7B Instruct: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BQwen2.5 7B Instruct
LMArena Vision1000—
GeoBench52%—

Instruction Following Not comparable

Llama 3.2 90B: —, Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BQwen2.5 7B Instruct
IFEval—74.1%

Writing & Preference Not comparable

Llama 3.2 90B: —, Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BQwen2.5 7B Instruct
WildBench—73.1%

Frequently asked questions

Is Llama 3.2 90B better than Qwen2.5 7B Instruct?

Qwen2.5 7B Instruct is the stronger model overall, scoring 29.0 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Qwen2.5 7B Instruct share?

5 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper