Model comparison
Llama2 70b Steerlm Chat vs Qwen1.5-72B
Llama2 70b Steerlm Chat and Qwen1.5-72B score almost the same on the Noometry Index (31.8 vs 30.8), so choose on price, context window or the category you care about most.
Last verified . 9 shared benchmarks.
Summary
- They share 9 benchmarks with published results for both. Llama2 70b Steerlm Chat scores higher in 0 categories and Qwen1.5-72B in 7 categories; 7 gaps are clear of the uncertainty.
- The widest gap is in writing & preference, where Qwen1.5-72B leads 37.3 to 31.6.
Side by side
| Llama2 70b Steerlm Chat | Qwen1.5-72B | |
|---|---|---|
| Provider | NVIDIA | Alibaba (Qwen) |
| Noometry Index | 31.8 | 30.8 |
| Released | — | 2024-02-04 |
| Weights | Open | Open |
| Context window | — | — |
| Max output | — | — |
| Input $ / M tokens | — | — |
| Output $ / M tokens | — | — |
| Results tracked | 9 | 22 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Qwen1.5-72B leads
Llama2 70b Steerlm Chat: 29.9 (#300), Qwen1.5-72B: 31.9 (#277)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-72B |
|---|---|---|
| LMArena Coding | 1025 | 1165 |
| BigCodeBench Instruct | — | 33.2% |
| BigCodeBench Complete | — | 40.3% |
| HumanEval+ | — | 59.1% |
| MBPP+ | — | 61.6% |
Reasoning Qwen1.5-72B leads
Llama2 70b Steerlm Chat: 20.0 (#246), Qwen1.5-72B: 22.2 (#203)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-72B |
|---|---|---|
| LMArena Hard Prompts | 1047 | 1148 |
Math Qwen1.5-72B leads
Llama2 70b Steerlm Chat: 31.3 (#226), Qwen1.5-72B: 33.2 (#205)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-72B |
|---|---|---|
| LMArena Math | 1072 | 1164 |
Knowledge Not comparable
Llama2 70b Steerlm Chat: —, Qwen1.5-72B: 11.5 (#300)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-72B |
|---|---|---|
| GPQA Diamond | — | 28.8% |
| LMArena Expert | — | 1136 |
Multilingual Qwen1.5-72B leads
Llama2 70b Steerlm Chat: 28.8 (#270), Qwen1.5-72B: 33.2 (#253)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-72B |
|---|---|---|
| LMArena Non-English | 1063 | 1135 |
| LMArena Chinese | — | 1186 |
| LMArena French | — | 1159 |
| LMArena German | — | 1084 |
| LMArena Japanese | — | 1061 |
| LMArena Korean | — | 1050 |
| LMArena Russian | — | 1104 |
| LMArena Spanish | — | 1110 |
Instruction Following Qwen1.5-72B leads
Llama2 70b Steerlm Chat: 54.2 (#279), Qwen1.5-72B: 59.3 (#256)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-72B |
|---|---|---|
| LMArena Instruction Following | 1060 | 1141 |
Long Context Qwen1.5-72B leads
Llama2 70b Steerlm Chat: 30.4 (#288), Qwen1.5-72B: 35.1 (#243)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-72B |
|---|---|---|
| LMArena Longer Query | 998 | 1157 |
Writing & Preference Qwen1.5-72B leads
Llama2 70b Steerlm Chat: 31.6 (#283), Qwen1.5-72B: 37.3 (#258)
| Benchmark | Llama2 70b Steerlm Chat | Qwen1.5-72B |
|---|---|---|
| LMArena Text | 1098 | 1166 |
| LMArena Creative Writing | 1091 | 1137 |
| LMArena Multi-Turn | 1058 | 1160 |
Frequently asked questions
Is Llama2 70b Steerlm Chat better than Qwen1.5-72B?
Llama2 70b Steerlm Chat and Qwen1.5-72B score almost the same on the Noometry Index (31.8 vs 30.8), so choose on price, context window or the category you care about most.
Is Llama2 70b Steerlm Chat or Qwen1.5-72B better for coding?
Qwen1.5-72B scores higher on coding benchmarks: 31.9 versus 29.9 in the Noometry coding category.
How many benchmarks do Llama2 70b Steerlm Chat and Qwen1.5-72B share?
9 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Qwen1.5-72B has 22.