Model comparison

Llama 3.2 90B vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 27.5 on the Noometry Index.

Last verified . 3 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Llama 3.2 90B scores higher in 0 categories and Qwen Max in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen Max leads 22.3 to 11.1.
  • The biggest single-benchmark swing is MATH Level 5: 39.4% for Llama 3.2 90B and 67.2% for Qwen Max.
  • Llama 3.2 90B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 90B and Qwen Max specifications
Llama 3.2 90BQwen Max
ProviderMetaAlibaba (Qwen)
Noometry Index27.534.7
Released2024-09-242024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked923

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkLlama 3.2 90BQwen Max
Aider Polyglot—21.8%
LMArena Coding—1288

Agentic & Tool Use Not comparable

Llama 3.2 90B: 30.0 (#80), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BQwen Max
BALROG27.3%—

Reasoning Qwen Max leads

Llama 3.2 90B: 21.7 (#217), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkLlama 3.2 90BQwen Max
EnigmaEval0.4%—
LMArena Hard Prompts—1269
Epoch Capabilities Index125.5—

Math Qwen Max leads

Llama 3.2 90B: 11.1 (#308), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkLlama 3.2 90BQwen Max
OTIS Mock AIME 2024-20252.6%16.1%
MATH Level 539.4%67.2%
LMArena Math—1275
FrontierMath (Feb 2025 set)—1%

Knowledge Qwen Max leads

Llama 3.2 90B: 21.7 (#274), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkLlama 3.2 90BQwen Max
GPQA Diamond41%56.1%
LMArena Expert—1248
MMLU80.3%—

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Qwen Max: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BQwen Max
LMArena Vision1000—
GeoBench52%—

Multilingual Not comparable

Llama 3.2 90B: —, Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkLlama 3.2 90BQwen Max
LMArena Non-English—1263
LMArena Chinese—1254
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Russian—1274
LMArena Spanish—1290

Instruction Following Not comparable

Llama 3.2 90B: —, Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BQwen Max
LMArena Instruction Following—1262

Long Context Not comparable

Llama 3.2 90B: —, Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkLlama 3.2 90BQwen Max
Fiction.LiveBench—66.7%
LMArena Longer Query—1288

Writing & Preference Not comparable

Llama 3.2 90B: —, Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BQwen Max
LMArena Text—1282
LMArena Creative Writing—1248
LMArena Multi-Turn—1277

Frequently asked questions

Is Llama 3.2 90B better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Qwen Max share?

3 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper