Model comparison

Llama 3.1 Nemotron Ultra 253b v1 vs Qwen3.5 35B-A3B

Qwen3.5 35B-A3B is the stronger model overall, scoring 42.0 to 36.7 on the Noometry Index.

Last verified . 10 shared benchmarks.

Qwen3.5 35B-A3B Alibaba (Qwen)

42.0

Rank #123 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.1 Nemotron Ultra 253b v1 scores higher in 2 categories and Qwen3.5 35B-A3B in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Qwen3.5 35B-A3B leads 50.0 to 43.1.

Side by side

Llama 3.1 Nemotron Ultra 253b v1 and Qwen3.5 35B-A3B specifications
Llama 3.1 Nemotron Ultra 253b v1Qwen3.5 35B-A3B
ProviderNVIDIAAlibaba (Qwen)
Noometry Index36.742.0
Released—2026-02-01
WeightsOpenOpen
Context window—262K
Max output—66K
Input $ / M tokens—$0.25
Output $ / M tokens—$2
Results tracked1128

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177), Qwen3.5 35B-A3B: 33.8 (#251)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 35B-A3B
LMArena Coding13121410
LMArena WebDev—1254
SciCode—29.3%

Agentic & Tool Use Not comparable

Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149), Qwen3.5 35B-A3B: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 35B-A3B
Berkeley Function Calling Leaderboard10%—

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134), Qwen3.5 35B-A3B: 24.6 (#161)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 35B-A3B
LMArena Hard Prompts13161400
CritPt—0.6%
Chess Puzzles—10%
DTBench—80%
LMCA—29.5%
Epoch Capabilities Index—142.52

Math Qwen3.5 35B-A3B leads

Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152), Qwen3.5 35B-A3B: 39.9 (#97)

Math benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 35B-A3B
LMArena Math13601404
MathArena Final-Answer Competitions—56%
OTIS Mock AIME 2024-2025—70%

Knowledge Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Qwen3.5 35B-A3B: 47.8 (#79)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 35B-A3B
GPQA Diamond—83.5%
Vectara Hallucination Rate—10.5%
LMArena Expert—1408

Multilingual Qwen3.5 35B-A3B leads

Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187), Qwen3.5 35B-A3B: 50.0 (#127)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 35B-A3B
LMArena Non-English12821378
LMArena Russian12841376
LMArena Chinese—1457
LMArena French—1412
LMArena German—1367
LMArena Japanese—1325
LMArena Korean—1356
LMArena Spanish—1392

Instruction Following Qwen3.5 35B-A3B leads

Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178), Qwen3.5 35B-A3B: 72.8 (#128)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 35B-A3B
LMArena Instruction Following13081379

Long Context Qwen3.5 35B-A3B leads

Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177), Qwen3.5 35B-A3B: 42.4 (#127)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 35B-A3B
LMArena Longer Query12991389

Writing & Preference Qwen3.5 35B-A3B leads

Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175), Qwen3.5 35B-A3B: 57.9 (#124)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.5 35B-A3B
LMArena Text13201395
LMArena Creative Writing13141346
LMArena Multi-Turn13171390

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253b v1 better than Qwen3.5 35B-A3B?

Qwen3.5 35B-A3B is the stronger model overall, scoring 42.0 to 36.7 on the Noometry Index.

Is Llama 3.1 Nemotron Ultra 253b v1 or Qwen3.5 35B-A3B better for coding?

Llama 3.1 Nemotron Ultra 253b v1 scores higher on coding benchmarks: 38.4 versus 33.8 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron Ultra 253b v1 and Qwen3.5 35B-A3B share?

10 benchmarks have published results for both models. Llama 3.1 Nemotron Ultra 253b v1 has 11 scored results on Noometry and Qwen3.5 35B-A3B has 28.

Related comparisons

Go deeper