Model comparison

Claude 3.5 Sonnet vs Llama 3.1 Nemotron Ultra 253b v1

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 34.6 on the Noometry Index.

Last verified . 10 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 5 categories and Llama 3.1 Nemotron Ultra 253b v1 in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama 3.1 Nemotron Ultra 253b v1 leads 37.5 to 19.2.
  • Llama 3.1 Nemotron Ultra 253b v1 has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Llama 3.1 Nemotron Ultra 253b v1 specifications
Claude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
ProviderAnthropicNVIDIA
Noometry Index34.636.7
Released2024-06-20—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked6011

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.5 Sonnet: 39.0 (#165), Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177)

Coding benchmarks
BenchmarkClaude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
LMArena Coding13421312
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 32.3 (#67), Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
Berkeley Function Calling Leaderboard—10%
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

Claude 3.5 Sonnet: 23.1 (#183), Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
LMArena Hard Prompts13051316
SimpleBench41.4%—
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
Epoch Capabilities Index133.55—
ForecastBench60.7—
LiveBench59%—

Math Llama 3.1 Nemotron Ultra 253b v1 leads

Claude 3.5 Sonnet: 19.2 (#288), Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152)

Math benchmarks
BenchmarkClaude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
LMArena Math13071360
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Not comparable

Claude 3.5 Sonnet: 28.6 (#245), Llama 3.1 Nemotron Ultra 253b v1: —

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
LMArena Expert1265—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Llama 3.1 Nemotron Ultra 253b v1: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Too close to call

Claude 3.5 Sonnet: 43.2 (#185), Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
LMArena Non-English12831282
LMArena Russian13061284
LMArena Chinese1272—
LMArena French1305—
LMArena German1297—
LMArena Japanese1234—
LMArena Korean1200—
LMArena Spanish1290—

Instruction Following Too close to call

Claude 3.5 Sonnet: 68.8 (#182), Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
LMArena Instruction Following12971308
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Too close to call

Claude 3.5 Sonnet: 39.9 (#167), Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
LMArena Longer Query13111299

Writing & Preference Too close to call

Claude 3.5 Sonnet: 52.9 (#164), Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetLlama 3.1 Nemotron Ultra 253b v1
LMArena Text12981320
LMArena Creative Writing12921314
LMArena Multi-Turn13261317
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Llama 3.1 Nemotron Ultra 253b v1?

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Llama 3.1 Nemotron Ultra 253b v1 better for coding?

They score almost the same on coding (39.0 vs 38.4); test both on your own repository before choosing.

How many benchmarks do Claude 3.5 Sonnet and Llama 3.1 Nemotron Ultra 253b v1 share?

10 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Llama 3.1 Nemotron Ultra 253b v1 has 11.

Related comparisons

Go deeper