Model comparison

Llama 3.2 90B vs Qwen Plus

Qwen Plus is the stronger model overall, scoring 37.1 to 27.5 on the Noometry Index.

Last verified . 3 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Llama 3.2 90B scores higher in 0 categories and Qwen Plus in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen Plus leads 23.3 to 11.1.
  • The biggest single-benchmark swing is MATH Level 5: 39.4% for Llama 3.2 90B and 65.3% for Qwen Plus.
  • Llama 3.2 90B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 90B and Qwen Plus specifications
Llama 3.2 90BQwen Plus
ProviderMetaAlibaba (Qwen)
Noometry Index27.537.1
Released2024-09-242024-01-25
WeightsOpenProprietary
Context window—1M
Max output—33K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.20
Results tracked920

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkLlama 3.2 90BQwen Plus
LMArena Coding—1328

Agentic & Tool Use Not comparable

Llama 3.2 90B: 30.0 (#80), Qwen Plus: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BQwen Plus
BALROG27.3%—

Reasoning Qwen Plus leads

Llama 3.2 90B: 21.7 (#217), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkLlama 3.2 90BQwen Plus
Kagi LLM Benchmark—63.3%
EnigmaEval0.4%—
LMArena Hard Prompts—1317
DTBench—81.1%
LMCA—24%
Epoch Capabilities Index125.5—

Math Qwen Plus leads

Llama 3.2 90B: 11.1 (#308), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkLlama 3.2 90BQwen Plus
OTIS Mock AIME 2024-20252.6%17.8%
MATH Level 539.4%65.3%
LMArena Math—1326
FrontierMath (Feb 2025 set)—1.7%

Knowledge Qwen Plus leads

Llama 3.2 90B: 21.7 (#274), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkLlama 3.2 90BQwen Plus
GPQA Diamond41%48.1%
LMArena Expert—1328
MMLU80.3%—

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Qwen Plus: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BQwen Plus
LMArena Vision1000—
GeoBench52%—

Multilingual Not comparable

Llama 3.2 90B: —, Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkLlama 3.2 90BQwen Plus
LMArena Non-English—1310
LMArena Chinese—1347
LMArena Japanese—1251
LMArena Russian—1323

Instruction Following Not comparable

Llama 3.2 90B: —, Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BQwen Plus
LMArena Instruction Following—1303

Long Context Not comparable

Llama 3.2 90B: —, Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkLlama 3.2 90BQwen Plus
LMArena Longer Query—1324

Writing & Preference Not comparable

Llama 3.2 90B: —, Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BQwen Plus
LMArena Text—1326
LMArena Creative Writing—1293
LMArena Multi-Turn—1336

Frequently asked questions

Is Llama 3.2 90B better than Qwen Plus?

Qwen Plus is the stronger model overall, scoring 37.1 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Qwen Plus share?

3 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper