Model comparison

Llama 3.2 90B vs Qwen1.5 4b Chat

Qwen1.5 4b Chat is the stronger model overall, scoring 28.8 to 27.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • The widest gap is in math, where Qwen1.5 4b Chat leads 30.4 to 11.1.

Side by side

Llama 3.2 90B and Qwen1.5 4b Chat specifications
Llama 3.2 90BQwen1.5 4b Chat
ProviderMetaAlibaba (Qwen)
Noometry Index27.528.8
Released2024-09-24—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked913

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkLlama 3.2 90BQwen1.5 4b Chat
LMArena Coding—999

Agentic & Tool Use Not comparable

Llama 3.2 90B: 30.0 (#80), Qwen1.5 4b Chat: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BQwen1.5 4b Chat
BALROG27.3%—

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkLlama 3.2 90BQwen1.5 4b Chat
EnigmaEval0.4%—
LMArena Hard Prompts—976
Epoch Capabilities Index125.5—

Math Qwen1.5 4b Chat leads

Llama 3.2 90B: 11.1 (#308), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkLlama 3.2 90BQwen1.5 4b Chat
OTIS Mock AIME 2024-20252.6%—
LMArena Math—1026
MATH Level 539.4%—

Knowledge Qwen1.5 4b Chat leads

Llama 3.2 90B: 21.7 (#274), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkLlama 3.2 90BQwen1.5 4b Chat
GPQA Diamond41%—
LMArena Expert—980
MMLU80.3%—

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Qwen1.5 4b Chat: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BQwen1.5 4b Chat
LMArena Vision1000—
GeoBench52%—

Multilingual Not comparable

Llama 3.2 90B: —, Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkLlama 3.2 90BQwen1.5 4b Chat
LMArena Non-English—979
LMArena Chinese—1024
LMArena German—902
LMArena Russian—952

Instruction Following Not comparable

Llama 3.2 90B: —, Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BQwen1.5 4b Chat
LMArena Instruction Following—978

Long Context Not comparable

Llama 3.2 90B: —, Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkLlama 3.2 90BQwen1.5 4b Chat
LMArena Longer Query—988

Writing & Preference Not comparable

Llama 3.2 90B: —, Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BQwen1.5 4b Chat
LMArena Text—997
LMArena Creative Writing—969
LMArena Multi-Turn—977

Frequently asked questions

Is Llama 3.2 90B better than Qwen1.5 4b Chat?

Qwen1.5 4b Chat is the stronger model overall, scoring 28.8 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Qwen1.5 4b Chat share?

0 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper