Model comparison

Llama2 70b Steerlm Chat vs Qwen2.5 32B Instruct

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 30.1 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Qwen2.5 32B Instruct Alibaba (Qwen)

30.1

Rank #297 Confirmed

Summary

  • The widest gap is in math, where Llama2 70b Steerlm Chat leads 31.3 to 16.2.

Side by side

Llama2 70b Steerlm Chat and Qwen2.5 32B Instruct specifications
Llama2 70b Steerlm ChatQwen2.5 32B Instruct
ProviderNVIDIAAlibaba (Qwen)
Noometry Index31.830.1
Released—2024-09
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.70
Output $ / M tokens—$2.80
Results tracked97

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 32B Instruct leads

Llama2 70b Steerlm Chat: 29.9 (#300), Qwen2.5 32B Instruct: 38.7 (#169)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen2.5 32B Instruct
BigCodeBench Instruct—45%
LMArena Coding1025—
BigCodeBench Complete—52.3%

Reasoning Too close to call

Llama2 70b Steerlm Chat: 20.0 (#246), Qwen2.5 32B Instruct: 19.2 (#266)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen2.5 32B Instruct
Chess Puzzles—0%
LMArena Hard Prompts1047—
Epoch Capabilities Index—128.52

Math Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 31.3 (#226), Qwen2.5 32B Instruct: 16.2 (#296)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen2.5 32B Instruct
OTIS Mock AIME 2024-2025—7.4%
LMArena Math1072—
MATH Level 5—56.1%

Knowledge Not comparable

Llama2 70b Steerlm Chat: —, Qwen2.5 32B Instruct: 24.9 (#266)

Knowledge benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen2.5 32B Instruct
GPQA Diamond—46.1%

Multilingual Not comparable

Llama2 70b Steerlm Chat: 28.8 (#270), Qwen2.5 32B Instruct: —

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen2.5 32B Instruct
LMArena Non-English1063—

Instruction Following Not comparable

Llama2 70b Steerlm Chat: 54.2 (#279), Qwen2.5 32B Instruct: —

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen2.5 32B Instruct
LMArena Instruction Following1060—

Long Context Not comparable

Llama2 70b Steerlm Chat: 30.4 (#288), Qwen2.5 32B Instruct: —

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen2.5 32B Instruct
LMArena Longer Query998—

Writing & Preference Not comparable

Llama2 70b Steerlm Chat: 31.6 (#283), Qwen2.5 32B Instruct: —

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen2.5 32B Instruct
LMArena Text1098—
LMArena Creative Writing1091—
LMArena Multi-Turn1058—

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Qwen2.5 32B Instruct?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 30.1 on the Noometry Index.

Is Llama2 70b Steerlm Chat or Qwen2.5 32B Instruct better for coding?

Qwen2.5 32B Instruct scores higher on coding benchmarks: 38.7 versus 29.9 in the Noometry coding category.

How many benchmarks do Llama2 70b Steerlm Chat and Qwen2.5 32B Instruct share?

0 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Qwen2.5 32B Instruct has 7.

Related comparisons

Go deeper