Model comparison

Llama 3.1 Nemotron Ultra 253b v1 vs Qwen3.8 27B

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 36.7 on the Noometry Index.

Last verified . 10 shared benchmarks.

Qwen3.8 27B Alibaba (Qwen)

46.0

Rank #68 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.1 Nemotron Ultra 253b v1 scores higher in 1 category and Qwen3.8 27B in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Qwen3.8 27B leads 32.9 to 15.7.

Side by side

Llama 3.1 Nemotron Ultra 253b v1 and Qwen3.8 27B specifications
Llama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
ProviderNVIDIAAlibaba (Qwen)
Noometry Index36.746.0
Released—2026-08-14
WeightsOpenOpen
Context window—262K
Max output—33K
Input $ / M tokens—$0.99
Output $ / M tokens—$1.49
Results tracked1131

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177), Qwen3.8 27B: 50.5 (#44)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
LMArena Coding13121482
LMArena WebDev—1593
SciCode—46.6%

Agentic & Tool Use Qwen3.8 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149), Qwen3.8 27B: 32.9 (#57)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
APEX-Agents—47.5%
Berkeley Function Calling Leaderboard10%—

Reasoning Qwen3.8 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134), Qwen3.8 27B: 41.0 (#54)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
LMArena Hard Prompts13161460
ARC-AGI-2—42.4%
NYT Connections (extended)—54.5%
ARC-AGI-1—87.5%
CritPt—5.4%
DTBench—88%
LMCA—41.4%
Surface Evolver Bench—45%
Epoch Capabilities Index—149.38

Math Too close to call

Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152), Qwen3.8 27B: 37.1 (#161)

Math benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
LMArena Math13601456
ProofBench—16%

Knowledge Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Qwen3.8 27B: 41.6 (#109)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
LMArena Expert—1482

Multimodal Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Qwen3.8 27B: 41.3 (#37)

Multimodal benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
LMArena Vision—1271

Multilingual Qwen3.8 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187), Qwen3.8 27B: 53.7 (#60)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
LMArena Non-English12821430
LMArena Russian12841415
LMArena Chinese—1504
LMArena French—1465
LMArena German—1438
LMArena Japanese—1384
LMArena Korean—1393
LMArena Spanish—1448

Instruction Following Qwen3.8 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178), Qwen3.8 27B: 75.8 (#53)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
LMArena Instruction Following13081439

Long Context Qwen3.8 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177), Qwen3.8 27B: 44.3 (#70)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
LMArena Longer Query12991450

Writing & Preference Qwen3.8 27B leads

Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175), Qwen3.8 27B: 65.8 (#43)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Qwen3.8 27B
LMArena Text13201441
LMArena Creative Writing13141384
LMArena Multi-Turn13171441
EQ-Bench Creative Writing—1671

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253b v1 better than Qwen3.8 27B?

Qwen3.8 27B is the stronger model overall, scoring 46.0 to 36.7 on the Noometry Index.

Is Llama 3.1 Nemotron Ultra 253b v1 or Qwen3.8 27B better for coding?

Qwen3.8 27B scores higher on coding benchmarks: 50.5 versus 38.4 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron Ultra 253b v1 and Qwen3.8 27B share?

10 benchmarks have published results for both models. Llama 3.1 Nemotron Ultra 253b v1 has 11 scored results on Noometry and Qwen3.8 27B has 31.

Related comparisons

Go deeper