Model comparison

Llama 3.1 Nemotron Ultra 253b v1 vs Qwen1.5 4b Chat

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 28.8 on the Noometry Index.

Last verified . 10 shared benchmarks.

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.1 Nemotron Ultra 253b v1 scores higher in 7 categories and Qwen1.5 4b Chat in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 3.1 Nemotron Ultra 253b v1 leads 52.2 to 23.8.

Side by side

Llama 3.1 Nemotron Ultra 253b v1 and Qwen1.5 4b Chat specifications
Llama 3.1 Nemotron Ultra 253b v1Qwen1.5 4b Chat
ProviderNVIDIAAlibaba (Qwen)
Noometry Index36.728.8
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1113

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen1.5 4b Chat
LMArena Coding1312999

Agentic & Tool Use Not comparable

Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149), Qwen1.5 4b Chat: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen1.5 4b Chat
Berkeley Function Calling Leaderboard10%—

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen1.5 4b Chat
LMArena Hard Prompts1316976

Math Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen1.5 4b Chat
LMArena Math13601026

Knowledge Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen1.5 4b Chat
LMArena Expert—980

Multilingual Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen1.5 4b Chat
LMArena Non-English1282979
LMArena Russian1284952
LMArena Chinese—1024
LMArena German—902

Instruction Following Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen1.5 4b Chat
LMArena Instruction Following1308978

Long Context Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen1.5 4b Chat
LMArena Longer Query1299988

Writing & Preference Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen1.5 4b Chat
LMArena Text1320997
LMArena Creative Writing1314969
LMArena Multi-Turn1317977

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253b v1 better than Qwen1.5 4b Chat?

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 28.8 on the Noometry Index.

Is Llama 3.1 Nemotron Ultra 253b v1 or Qwen1.5 4b Chat better for coding?

Llama 3.1 Nemotron Ultra 253b v1 scores higher on coding benchmarks: 38.4 versus 29.1 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron Ultra 253b v1 and Qwen1.5 4b Chat share?

10 benchmarks have published results for both models. Llama 3.1 Nemotron Ultra 253b v1 has 11 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper