Model comparison
Llama2 70b Steerlm Chat vs Qwen1.5-14B
Llama2 70b Steerlm Chat and Qwen1.5-14B score almost the same on the Noometry Index (31.8 vs 32.7), so choose on price, context window or the category you care about most.
Last verified . 9 shared benchmarks.
Summary
- They share 9 benchmarks with published results for both. Llama2 70b Steerlm Chat scores higher in 0 categories and Qwen1.5-14B in 7 categories; 7 gaps are clear of the uncertainty.
- The widest gap is in long context, where Qwen1.5-14B leads 33.7 to 30.4.
Side by side
| Llama2 70b Steerlm Chat | Qwen1.5-14B | |
|---|---|---|
| Provider | NVIDIA | Alibaba (Qwen) |
| Noometry Index | 31.8 | 32.7 |
| Released | — | 2024-02-04 |
| Weights | Open | Open |
| Context window | — | — |
| Max output | — | — |
| Input $ / M tokens | — | — |
| Output $ / M tokens | — | — |
| Results tracked | 9 | 17 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Qwen1.5-14B leads
Llama2 70b Steerlm Chat: 29.9 (#300), Qwen1.5-14B: 33.1 (#263)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-14B |
|---|---|---|
| LMArena Coding | 1025 | 1138 |
Reasoning Qwen1.5-14B leads
Llama2 70b Steerlm Chat: 20.0 (#246), Qwen1.5-14B: 21.4 (#223)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-14B |
|---|---|---|
| LMArena Hard Prompts | 1047 | 1113 |
Math Qwen1.5-14B leads
Llama2 70b Steerlm Chat: 31.3 (#226), Qwen1.5-14B: 32.4 (#215)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-14B |
|---|---|---|
| LMArena Math | 1072 | 1125 |
Knowledge Not comparable
Llama2 70b Steerlm Chat: —, Qwen1.5-14B: 29.8 (#232)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-14B |
|---|---|---|
| LMArena Expert | — | 1094 |
| MMLU | — | 68.6% |
Multilingual Qwen1.5-14B leads
Llama2 70b Steerlm Chat: 28.8 (#270), Qwen1.5-14B: 30.7 (#262)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-14B |
|---|---|---|
| LMArena Non-English | 1063 | 1095 |
| LMArena Chinese | — | 1147 |
| LMArena French | — | 1116 |
| LMArena German | — | 1043 |
| LMArena Japanese | — | 1019 |
| LMArena Russian | — | 1046 |
| LMArena Spanish | — | 1085 |
Instruction Following Qwen1.5-14B leads
Llama2 70b Steerlm Chat: 54.2 (#279), Qwen1.5-14B: 56.8 (#271)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-14B |
|---|---|---|
| LMArena Instruction Following | 1060 | 1102 |
Long Context Qwen1.5-14B leads
Llama2 70b Steerlm Chat: 30.4 (#288), Qwen1.5-14B: 33.7 (#257)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-14B |
|---|---|---|
| LMArena Longer Query | 998 | 1113 |
Writing & Preference Qwen1.5-14B leads
Llama2 70b Steerlm Chat: 31.6 (#283), Qwen1.5-14B: 33.6 (#276)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-14B |
|---|---|---|
| LMArena Text | 1098 | 1128 |
| LMArena Creative Writing | 1091 | 1091 |
| LMArena Multi-Turn | 1058 | 1110 |
Frequently asked questions
Is Llama2 70b Steerlm Chat better than Qwen1.5-14B?
Llama2 70b Steerlm Chat and Qwen1.5-14B score almost the same on the Noometry Index (31.8 vs 32.7), so choose on price, context window or the category you care about most.
Is Llama2 70b Steerlm Chat or Qwen1.5-14B better for coding?
Qwen1.5-14B scores higher on coding benchmarks: 33.1 versus 29.9 in the Noometry coding category.
How many benchmarks do Llama2 70b Steerlm Chat and Qwen1.5-14B share?
9 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Qwen1.5-14B has 17.