Model comparison

Llama 3.1 Nemotron 70b Instruct vs Qwen Max

Llama 3.1 Nemotron 70b Instruct is the stronger model overall, scoring 37.6 to 34.7 on the Noometry Index.

Last verified . 12 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Llama 3.1 Nemotron 70b Instruct scores higher in 4 categories and Qwen Max in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama 3.1 Nemotron 70b Instruct leads 35.5 to 22.3.
  • Llama 3.1 Nemotron 70b Instruct has downloadable open weights; the other is API-only.

Side by side

Llama 3.1 Nemotron 70b Instruct and Qwen Max specifications
Llama 3.1 Nemotron 70b InstructQwen Max
ProviderNVIDIAAlibaba (Qwen)
Noometry Index37.634.7
Released2024-12-182024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked1423

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron 70b Instruct leads

Llama 3.1 Nemotron 70b Instruct: 35.9 (#216), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructQwen Max
LMArena Coding12721288
Aider Polyglot—21.8%
BigCodeBench Instruct38.7%—
BigCodeBench Complete48.2%—

Reasoning Too close to call

Llama 3.1 Nemotron 70b Instruct: 25.0 (#152), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructQwen Max
LMArena Hard Prompts12661269

Math Llama 3.1 Nemotron 70b Instruct leads

Llama 3.1 Nemotron 70b Instruct: 35.5 (#182), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructQwen Max
LMArena Math12711275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Llama 3.1 Nemotron 70b Instruct leads

Llama 3.1 Nemotron 70b Instruct: 34.1 (#199), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructQwen Max
LMArena Expert12421248
GPQA Diamond—56.1%

Multilingual Qwen Max leads

Llama 3.1 Nemotron 70b Instruct: 40.5 (#217), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructQwen Max
LMArena Non-English12451263
LMArena Chinese12631254
LMArena Russian12271274
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Spanish—1290

Instruction Following Too close to call

Llama 3.1 Nemotron 70b Instruct: 65.9 (#213), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructQwen Max
LMArena Instruction Following12521262

Long Context Qwen Max leads

Llama 3.1 Nemotron 70b Instruct: 37.6 (#215), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructQwen Max
LMArena Longer Query12381288
Fiction.LiveBench—66.7%

Writing & Preference Too close to call

Llama 3.1 Nemotron 70b Instruct: 48.4 (#203), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructQwen Max
LMArena Text12831282
LMArena Creative Writing12691248
LMArena Multi-Turn12751277

Frequently asked questions

Is Llama 3.1 Nemotron 70b Instruct better than Qwen Max?

Llama 3.1 Nemotron 70b Instruct is the stronger model overall, scoring 37.6 to 34.7 on the Noometry Index.

Is Llama 3.1 Nemotron 70b Instruct or Qwen Max better for coding?

Llama 3.1 Nemotron 70b Instruct scores higher on coding benchmarks: 35.9 versus 30.7 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron 70b Instruct and Qwen Max share?

12 benchmarks have published results for both models. Llama 3.1 Nemotron 70b Instruct has 14 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper