Model comparison

Claude 3 Sonnet vs Nemotron 3 Super

Nemotron 3 Super is the stronger model overall, scoring 40.1 to 29.0 on the Noometry Index.

Last verified . 16 shared benchmarks.

Claude 3 Sonnet Anthropic

29.0

Rank #319 Confirmed

Nemotron 3 Super NVIDIA

40.1

Rank #153 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Claude 3 Sonnet scores higher in 1 category and Nemotron 3 Super in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Nemotron 3 Super leads 39.6 to 10.7.
  • The biggest single-benchmark swing is WeirdML: 10.2% for Claude 3 Sonnet and 38% for Nemotron 3 Super.
  • Nemotron 3 Super has downloadable open weights; the other is API-only.

Side by side

Claude 3 Sonnet and Nemotron 3 Super specifications
Claude 3 SonnetNemotron 3 Super
ProviderAnthropicNVIDIA
Noometry Index29.040.1
Released2024-02-292026-03-11
WeightsProprietaryOpen
Context window—262K
Max output—131K
Input $ / M tokens—$0.08
Output $ / M tokens—$0.45
Results tracked3021

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nemotron 3 Super leads

Claude 3 Sonnet: 29.6 (#302), Nemotron 3 Super: 39.4 (#158)

Coding benchmarks
BenchmarkClaude 3 SonnetNemotron 3 Super
WeirdML10.2%38%
LMArena Coding12231403
SciCode—36%
BigCodeBench Instruct42.7%—
BigCodeBench Complete53.8%—
ALE-Bench—213.9
HumanEval+64%—
MBPP+69.3%—

Reasoning Claude 3 Sonnet leads

Claude 3 Sonnet: 20.5 (#237), Nemotron 3 Super: 18.7 (#276)

Reasoning benchmarks
BenchmarkClaude 3 SonnetNemotron 3 Super
LMArena Hard Prompts11971386
NYT Connections (extended)—15.4%
CritPt—3.1%
DTBench53.6%—
Epoch Capabilities Index120.7—
WinoGrande75.1%—

Math Nemotron 3 Super leads

Claude 3 Sonnet: 10.7 (#310), Nemotron 3 Super: 39.6 (#101)

Math benchmarks
BenchmarkClaude 3 SonnetNemotron 3 Super
LMArena Math12131380
MathArena Final-Answer Competitions—60.4%
OTIS Mock AIME 2024-20252.5%—
MATH Level 518.2%—

Knowledge Nemotron 3 Super leads

Claude 3 Sonnet: 21.1 (#276), Nemotron 3 Super: 38.9 (#139)

Knowledge benchmarks
BenchmarkClaude 3 SonnetNemotron 3 Super
LMArena Expert11731398
GPQA Diamond40.6%—
MMLU75.9%—

Multimodal Not comparable

Claude 3 Sonnet: 25.2 (#125), Nemotron 3 Super: —

Multimodal benchmarks
BenchmarkClaude 3 SonnetNemotron 3 Super
LMArena Vision984—

Multilingual Nemotron 3 Super leads

Claude 3 Sonnet: 37.8 (#234), Nemotron 3 Super: 48.4 (#144)

Multilingual benchmarks
BenchmarkClaude 3 SonnetNemotron 3 Super
LMArena Non-English12051355
LMArena Chinese11891432
LMArena French12291405
LMArena German12041341
LMArena Russian12271336
LMArena Spanish12041417
LMArena Japanese1131—
LMArena Korean1128—

Instruction Following Nemotron 3 Super leads

Claude 3 Sonnet: 62.8 (#235), Nemotron 3 Super: 71.2 (#154)

Instruction Following benchmarks
BenchmarkClaude 3 SonnetNemotron 3 Super
LMArena Instruction Following11991347

Long Context Nemotron 3 Super leads

Claude 3 Sonnet: 36.7 (#228), Nemotron 3 Super: 41.5 (#139)

Long Context benchmarks
BenchmarkClaude 3 SonnetNemotron 3 Super
LMArena Longer Query12111362

Writing & Preference Nemotron 3 Super leads

Claude 3 Sonnet: 42.1 (#238), Nemotron 3 Super: 55.9 (#140)

Writing & Preference benchmarks
BenchmarkClaude 3 SonnetNemotron 3 Super
LMArena Text12181378
LMArena Creative Writing11861317
LMArena Multi-Turn12271369

Frequently asked questions

Is Claude 3 Sonnet better than Nemotron 3 Super?

Nemotron 3 Super is the stronger model overall, scoring 40.1 to 29.0 on the Noometry Index.

Is Claude 3 Sonnet or Nemotron 3 Super better for coding?

Nemotron 3 Super scores higher on coding benchmarks: 39.4 versus 29.6 in the Noometry coding category.

How many benchmarks do Claude 3 Sonnet and Nemotron 3 Super share?

16 benchmarks have published results for both models. Claude 3 Sonnet has 30 scored results on Noometry and Nemotron 3 Super has 21.

Related comparisons

Go deeper