Model comparison

Llama2 70b Steerlm Chat vs Qwen1.5-7B

Llama2 70b Steerlm Chat and Qwen1.5-7B score almost the same on the Noometry Index (31.8 vs 31.4), so choose on price, context window or the category you care about most.

Last verified . 9 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Qwen1.5-7B Alibaba (Qwen)

31.4

Rank #273 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama2 70b Steerlm Chat scores higher in 3 categories and Qwen1.5-7B in 4 categories; 3 gaps are clear of the uncertainty.

Side by side

Llama2 70b Steerlm Chat and Qwen1.5-7B specifications
Llama2 70b Steerlm ChatQwen1.5-7B
ProviderNVIDIAAlibaba (Qwen)
Noometry Index31.831.4
Released—2024-02-04
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked913

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5-7B leads

Llama2 70b Steerlm Chat: 29.9 (#300), Qwen1.5-7B: 32.2 (#276)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-7B
LMArena Coding10251107

Reasoning Too close to call

Llama2 70b Steerlm Chat: 20.0 (#246), Qwen1.5-7B: 20.4 (#240)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-7B
LMArena Hard Prompts10471065

Math Too close to call

Llama2 70b Steerlm Chat: 31.3 (#226), Qwen1.5-7B: 31.4 (#224)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-7B
LMArena Math10721080

Knowledge Not comparable

Llama2 70b Steerlm Chat: —, Qwen1.5-7B: 28.7 (#243)

Knowledge benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-7B
LMArena Expert—1055
MMLU—62.6%

Multilingual Too close to call

Llama2 70b Steerlm Chat: 28.8 (#270), Qwen1.5-7B: 28.5 (#271)

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-7B
LMArena Non-English10631058
LMArena Chinese—1141
LMArena Russian—1006

Instruction Following Too close to call

Llama2 70b Steerlm Chat: 54.2 (#279), Qwen1.5-7B: 54.1 (#281)

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-7B
LMArena Instruction Following10601058

Long Context Qwen1.5-7B leads

Llama2 70b Steerlm Chat: 30.4 (#288), Qwen1.5-7B: 33.1 (#266)

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-7B
LMArena Longer Query9981090

Writing & Preference Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 31.6 (#283), Qwen1.5-7B: 29.6 (#293)

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-7B
LMArena Text10981083
LMArena Creative Writing10911035
LMArena Multi-Turn10581062

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Qwen1.5-7B?

Llama2 70b Steerlm Chat and Qwen1.5-7B score almost the same on the Noometry Index (31.8 vs 31.4), so choose on price, context window or the category you care about most.

Is Llama2 70b Steerlm Chat or Qwen1.5-7B better for coding?

Qwen1.5-7B scores higher on coding benchmarks: 32.2 versus 29.9 in the Noometry coding category.

How many benchmarks do Llama2 70b Steerlm Chat and Qwen1.5-7B share?

9 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Qwen1.5-7B has 13.

Related comparisons

Go deeper