Model comparison

Llama 3.2 90B vs Llama2 70b Steerlm Chat

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 27.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • The widest gap is in math, where Llama2 70b Steerlm Chat leads 31.3 to 11.1.

Side by side

Llama 3.2 90B and Llama2 70b Steerlm Chat specifications
Llama 3.2 90BLlama2 70b Steerlm Chat
ProviderMetaNVIDIA
Noometry Index27.531.8
Released2024-09-24—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked99

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkLlama 3.2 90BLlama2 70b Steerlm Chat
LMArena Coding—1025

Agentic & Tool Use Not comparable

Llama 3.2 90B: 30.0 (#80), Llama2 70b Steerlm Chat: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BLlama2 70b Steerlm Chat
BALROG27.3%—

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkLlama 3.2 90BLlama2 70b Steerlm Chat
EnigmaEval0.4%—
LMArena Hard Prompts—1047
Epoch Capabilities Index125.5—

Math Llama2 70b Steerlm Chat leads

Llama 3.2 90B: 11.1 (#308), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkLlama 3.2 90BLlama2 70b Steerlm Chat
OTIS Mock AIME 2024-20252.6%—
LMArena Math—1072
MATH Level 539.4%—

Knowledge Not comparable

Llama 3.2 90B: 21.7 (#274), Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkLlama 3.2 90BLlama2 70b Steerlm Chat
GPQA Diamond41%—
MMLU80.3%—

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Llama2 70b Steerlm Chat: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BLlama2 70b Steerlm Chat
LMArena Vision1000—
GeoBench52%—

Multilingual Not comparable

Llama 3.2 90B: —, Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkLlama 3.2 90BLlama2 70b Steerlm Chat
LMArena Non-English—1063

Instruction Following Not comparable

Llama 3.2 90B: —, Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BLlama2 70b Steerlm Chat
LMArena Instruction Following—1060

Long Context Not comparable

Llama 3.2 90B: —, Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkLlama 3.2 90BLlama2 70b Steerlm Chat
LMArena Longer Query—998

Writing & Preference Not comparable

Llama 3.2 90B: —, Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BLlama2 70b Steerlm Chat
LMArena Text—1098
LMArena Creative Writing—1091
LMArena Multi-Turn—1058

Frequently asked questions

Is Llama 3.2 90B better than Llama2 70b Steerlm Chat?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Llama2 70b Steerlm Chat share?

0 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper