Model comparison

Llama 3.1 Nemotron Ultra 253b v1 vs Qwen3.5 27B

Qwen3.5 27B is the stronger model overall, scoring 41.9 to 36.7 on the Noometry Index.

Last verified . 10 shared benchmarks.

Qwen3.5 27B Alibaba (Qwen)

41.9

Rank #127 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.1 Nemotron Ultra 253b v1 scores higher in 0 categories and Qwen3.5 27B in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Qwen3.5 27B leads 50.8 to 43.1.

Side by side

Llama 3.1 Nemotron Ultra 253b v1 and Qwen3.5 27B specifications
Llama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
ProviderNVIDIAAlibaba (Qwen)
Noometry Index36.741.9
Released—2026-02-23
WeightsOpenOpen
Context window—262K
Max output—66K
Input $ / M tokens—$0.30
Output $ / M tokens—$2.40
Results tracked1128

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177), Qwen3.5 27B: 38.9 (#168)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
LMArena Coding13121427
LMArena WebDev—1358
WeirdML—39.5%
ALE-Bench—349.45

Agentic & Tool Use Not comparable

Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149), Qwen3.5 27B: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
Berkeley Function Calling Leaderboard10%—
Vending-Bench 2—201.98

Reasoning Qwen3.5 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134), Qwen3.5 27B: 27.5 (#117)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
LMArena Hard Prompts13161414
NYT Connections (extended)—47.9%
Thematic Generalization—45.5%
DTBench—82.4%
LMCA—34%

Math Qwen3.5 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152), Qwen3.5 27B: 38.8 (#127)

Math benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
LMArena Math13601429
MathArena Final-Answer Competitions—56.7%

Knowledge Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Qwen3.5 27B: 38.0 (#150)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
Vectara Hallucination Rate—12.1%
LMArena Expert—1428

Multimodal Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Qwen3.5 27B: 39.4 (#59)

Multimodal benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
LMArena Vision—1241

Multilingual Qwen3.5 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187), Qwen3.5 27B: 50.8 (#115)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
LMArena Non-English12821390
LMArena Russian12841390
LMArena Chinese—1478
LMArena French—1410
LMArena German—1393
LMArena Japanese—1345
LMArena Korean—1358
LMArena Spanish—1407

Instruction Following Qwen3.5 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178), Qwen3.5 27B: 73.5 (#119)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
LMArena Instruction Following13081393

Long Context Qwen3.5 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177), Qwen3.5 27B: 43.1 (#106)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
LMArena Longer Query12991413

Writing & Preference Qwen3.5 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175), Qwen3.5 27B: 59.3 (#111)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 27B
LMArena Text13201409
LMArena Creative Writing13141362
LMArena Multi-Turn13171410

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253b v1 better than Qwen3.5 27B?

Qwen3.5 27B is the stronger model overall, scoring 41.9 to 36.7 on the Noometry Index.

Is Llama 3.1 Nemotron Ultra 253b v1 or Qwen3.5 27B better for coding?

They score almost the same on coding (38.4 vs 38.9); test both on your own repository before choosing.

How many benchmarks do Llama 3.1 Nemotron Ultra 253b v1 and Qwen3.5 27B share?

10 benchmarks have published results for both models. Llama 3.1 Nemotron Ultra 253b v1 has 11 scored results on Noometry and Qwen3.5 27B has 28.

Related comparisons

Go deeper