Model comparison

Nvidia Llama 3.3 Nemotron Super 49b v1.5 vs Qwen2.5 7B Instruct

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 29.0 on the Noometry Index.

Last verified . 0 shared benchmarks.

Summary

  • The widest gap is in math, where Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads 38.2 to 12.6.
  • Qwen2.5 7B Instruct is cheaper at $0.17 / $0.70 per million input/output tokens, against $0.40 / $0.40 for Nvidia Llama 3.3 Nemotron Super 49b v1.5.

Side by side

Nvidia Llama 3.3 Nemotron Super 49b v1.5 and Qwen2.5 7B Instruct specifications
Nvidia Llama 3.3 Nemotron Super 49b v1.5Qwen2.5 7B Instruct
ProviderNVIDIAAlibaba (Qwen)
Noometry Index40.329.0
Released2025-07-252024-09
WeightsOpenOpen
Context window131K131K
Max output131K8K
Input $ / M tokens$0.40$0.17
Output $ / M tokens$0.40$0.70
Results tracked1215

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 39.8 (#154), Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Qwen2.5 7B Instruct
BigCodeBench Instruct—37.6%
LMArena Coding1355—
BigCodeBench Complete—46.1%

Agentic & Tool Use Not comparable

Nvidia Llama 3.3 Nemotron Super 49b v1.5: —, Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Qwen2.5 7B Instruct
BALROG—7.8%

Reasoning Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 26.8 (#128), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Qwen2.5 7B Instruct
Chess Puzzles—0%
LMArena Hard Prompts1336—
DTBench—47.7%
LMCA—6.4%
Epoch Capabilities Index—118.51

Math Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 38.2 (#141), Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Qwen2.5 7B Instruct
OTIS Mock AIME 2024-2025—2.5%
Omni-MATH—29.4%
LMArena Math1392—

Knowledge Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 36.7 (#165), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Qwen2.5 7B Instruct
GPQA Diamond—35.5%
MMLU-Pro—53.9%
GPQA (HELM)—34.1%
LMArena Expert1330—
MMLU—72.9%

Multilingual Not comparable

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 45.5 (#168), Qwen2.5 7B Instruct: —

Multilingual benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Qwen2.5 7B Instruct
LMArena Non-English1316—
LMArena Japanese1300—
LMArena Russian1332—

Instruction Following Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 68.6 (#188), Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Qwen2.5 7B Instruct
IFEval—74.1%
LMArena Instruction Following1299—

Long Context Not comparable

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 40.0 (#164), Qwen2.5 7B Instruct: —

Long Context benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Qwen2.5 7B Instruct
LMArena Longer Query1315—

Writing & Preference Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 53.1 (#159), Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Qwen2.5 7B Instruct
LMArena Text1338—
LMArena Creative Writing1307—
WildBench—73.1%
LMArena Multi-Turn1334—

Frequently asked questions

Is Nvidia Llama 3.3 Nemotron Super 49b v1.5 better than Qwen2.5 7B Instruct?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 29.0 on the Noometry Index.

Which is cheaper, Nvidia Llama 3.3 Nemotron Super 49b v1.5 or Qwen2.5 7B Instruct?

Qwen2.5 7B Instruct is cheaper. It lists at $0.17 per million input tokens and $0.70 per million output tokens; Nvidia Llama 3.3 Nemotron Super 49b v1.5 lists at $0.40 and $0.40.

Is Nvidia Llama 3.3 Nemotron Super 49b v1.5 or Qwen2.5 7B Instruct better for coding?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 scores higher on coding benchmarks: 39.8 versus 36.5 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Nvidia Llama 3.3 Nemotron Super 49b v1.5 and Qwen2.5 7B Instruct share?

0 benchmarks have published results for both models. Nvidia Llama 3.3 Nemotron Super 49b v1.5 has 12 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper