Model comparison

Llama 3.3 Nemotron 49b Super v1 vs Qwen1.5-7B

Llama 3.3 Nemotron 49b Super v1 is the stronger model overall, scoring 40.1 to 31.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

Qwen1.5-7B Alibaba (Qwen)

31.4

Rank #273 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.3 Nemotron 49b Super v1 scores higher in 6 categories and Qwen1.5-7B in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 3.3 Nemotron 49b Super v1 leads 50.8 to 29.6.

Side by side

Llama 3.3 Nemotron 49b Super v1 and Qwen1.5-7B specifications
Llama 3.3 Nemotron 49b Super v1Qwen1.5-7B
ProviderNVIDIAAlibaba (Qwen)
Noometry Index40.131.4
Released—2024-02-04
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1013

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 37.9 (#186), Qwen1.5-7B: 32.2 (#276)

Coding benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Qwen1.5-7B
LMArena Coding12961107

Reasoning Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 26.2 (#135), Qwen1.5-7B: 20.4 (#240)

Reasoning benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Qwen1.5-7B
LMArena Hard Prompts13111065

Math Not comparable

Llama 3.3 Nemotron 49b Super v1: —, Qwen1.5-7B: 31.4 (#224)

Math benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Qwen1.5-7B
LMArena Math—1080

Knowledge Not comparable

Llama 3.3 Nemotron 49b Super v1: —, Qwen1.5-7B: 28.7 (#243)

Knowledge benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Qwen1.5-7B
LMArena Expert—1055
MMLU—62.6%

Multilingual Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 41.1 (#211), Qwen1.5-7B: 28.5 (#271)

Multilingual benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Qwen1.5-7B
LMArena Non-English12531058
LMArena Chinese12771141
LMArena Russian12691006

Instruction Following Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 68.3 (#189), Qwen1.5-7B: 54.1 (#281)

Instruction Following benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Qwen1.5-7B
LMArena Instruction Following12931058

Long Context Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 39.5 (#176), Qwen1.5-7B: 33.1 (#266)

Long Context benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Qwen1.5-7B
LMArena Longer Query12991090

Writing & Preference Llama 3.3 Nemotron 49b Super v1 leads

Llama 3.3 Nemotron 49b Super v1: 50.8 (#179), Qwen1.5-7B: 29.6 (#293)

Writing & Preference benchmarks
BenchmarkLlama 3.3 Nemotron 49b Super v1Qwen1.5-7B
LMArena Text13081083
LMArena Creative Writing12881035
LMArena Multi-Turn13151062

Frequently asked questions

Is Llama 3.3 Nemotron 49b Super v1 better than Qwen1.5-7B?

Llama 3.3 Nemotron 49b Super v1 is the stronger model overall, scoring 40.1 to 31.4 on the Noometry Index.

Is Llama 3.3 Nemotron 49b Super v1 or Qwen1.5-7B better for coding?

Llama 3.3 Nemotron 49b Super v1 scores higher on coding benchmarks: 37.9 versus 32.2 in the Noometry coding category.

How many benchmarks do Llama 3.3 Nemotron 49b Super v1 and Qwen1.5-7B share?

10 benchmarks have published results for both models. Llama 3.3 Nemotron 49b Super v1 has 10 scored results on Noometry and Qwen1.5-7B has 13.

Related comparisons

Go deeper