Model comparison

Llama2 70b Steerlm Chat vs Qwen1.5-14B

Llama2 70b Steerlm Chat and Qwen1.5-14B score almost the same on the Noometry Index (31.8 vs 32.7), so choose on price, context window or the category you care about most.

Last verified . 9 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Qwen1.5-14B Alibaba (Qwen)

32.7

Rank #253 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama2 70b Steerlm Chat scores higher in 0 categories and Qwen1.5-14B in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Qwen1.5-14B leads 33.7 to 30.4.

Side by side

Llama2 70b Steerlm Chat and Qwen1.5-14B specifications
Llama2 70b Steerlm ChatQwen1.5-14B
ProviderNVIDIAAlibaba (Qwen)
Noometry Index31.832.7
Released—2024-02-04
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked917

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5-14B leads

Llama2 70b Steerlm Chat: 29.9 (#300), Qwen1.5-14B: 33.1 (#263)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-14B
LMArena Coding10251138

Reasoning Qwen1.5-14B leads

Llama2 70b Steerlm Chat: 20.0 (#246), Qwen1.5-14B: 21.4 (#223)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-14B
LMArena Hard Prompts10471113

Math Qwen1.5-14B leads

Llama2 70b Steerlm Chat: 31.3 (#226), Qwen1.5-14B: 32.4 (#215)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-14B
LMArena Math10721125

Knowledge Not comparable

Llama2 70b Steerlm Chat: —, Qwen1.5-14B: 29.8 (#232)

Knowledge benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-14B
LMArena Expert—1094
MMLU—68.6%

Multilingual Qwen1.5-14B leads

Llama2 70b Steerlm Chat: 28.8 (#270), Qwen1.5-14B: 30.7 (#262)

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-14B
LMArena Non-English10631095
LMArena Chinese—1147
LMArena French—1116
LMArena German—1043
LMArena Japanese—1019
LMArena Russian—1046
LMArena Spanish—1085

Instruction Following Qwen1.5-14B leads

Llama2 70b Steerlm Chat: 54.2 (#279), Qwen1.5-14B: 56.8 (#271)

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-14B
LMArena Instruction Following10601102

Long Context Qwen1.5-14B leads

Llama2 70b Steerlm Chat: 30.4 (#288), Qwen1.5-14B: 33.7 (#257)

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-14B
LMArena Longer Query9981113

Writing & Preference Qwen1.5-14B leads

Llama2 70b Steerlm Chat: 31.6 (#283), Qwen1.5-14B: 33.6 (#276)

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatQwen1.5-14B
LMArena Text10981128
LMArena Creative Writing10911091
LMArena Multi-Turn10581110

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Qwen1.5-14B?

Llama2 70b Steerlm Chat and Qwen1.5-14B score almost the same on the Noometry Index (31.8 vs 32.7), so choose on price, context window or the category you care about most.

Is Llama2 70b Steerlm Chat or Qwen1.5-14B better for coding?

Qwen1.5-14B scores higher on coding benchmarks: 33.1 versus 29.9 in the Noometry coding category.

How many benchmarks do Llama2 70b Steerlm Chat and Qwen1.5-14B share?

9 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Qwen1.5-14B has 17.

Related comparisons

Go deeper