Model comparison

Claude Opus 4.1 vs Nemotron 3 Ultra

Nemotron 3 Ultra is the stronger model overall, scoring 42.5 to 41.0 on the Noometry Index.

Last verified . 24 shared benchmarks.

Claude Opus 4.1 Anthropic

41.0

Rank #142 Confirmed

Nemotron 3 Ultra NVIDIA

42.5

Rank #113 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Claude Opus 4.1 scores higher in 5 categories and Nemotron 3 Ultra in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Nemotron 3 Ultra leads 35.3 to 22.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 68.9% for Claude Opus 4.1 and 86.7% for Nemotron 3 Ultra.
  • Nemotron 3 Ultra is cheaper at $0.50 / $2.20 per million input/output tokens, against $15 / $75 for Claude Opus 4.1.
  • Nemotron 3 Ultra accepts more context: 262K tokens versus 200K.
  • Nemotron 3 Ultra has downloadable open weights; the other is API-only.

Side by side

Claude Opus 4.1 and Nemotron 3 Ultra specifications
Claude Opus 4.1Nemotron 3 Ultra
ProviderAnthropicNVIDIA
Noometry Index41.042.5
Released2025-08-052026-06-04
WeightsProprietaryOpen
Context window200K262K
Max output32K128K
Input $ / M tokens$15$0.50
Output $ / M tokens$75$2.20
Results tracked4830

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.1 leads

Claude Opus 4.1: 44.4 (#73), Nemotron 3 Ultra: 38.1 (#182)

Coding benchmarks
BenchmarkClaude Opus 4.1Nemotron 3 Ultra
WeirdML45.9%43.5%
LMArena Coding14791468
SWE-bench Verified73.3%—
FrontierCode—13.6%
LMArena WebDev1390—
SciCode—40.3%
ALE-Bench674.77—
AlgoTune1.34—

Agentic & Tool Use Claude Opus 4.1 leads

Claude Opus 4.1: 35.0 (#41), Nemotron 3 Ultra: 23.4 (#126)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.1Nemotron 3 Ultra
Terminal-Bench38%—
APEX-Agents—22.7%
GDPval43.6%—
Cybench42%—
DeepResearch Bench48.3%—
LMArena Search1148—
METR Time Horizons66.8%—

Reasoning Claude Opus 4.1 leads

Claude Opus 4.1: 32.2 (#76), Nemotron 3 Ultra: 31.0 (#82)

Reasoning benchmarks
BenchmarkClaude Opus 4.1Nemotron 3 Ultra
Chess Puzzles7%12%
LMArena Hard Prompts14431452
Mystery Game Puzzles21%20%
DTBench80%90.1%
LMCA37.1%36.9%
Epoch Capabilities Index144.12146.17
SimpleBench60%—
CritPt—3.1%
EnigmaEval7.2%—
EBR-Bench7.9%—
ForecastBench62—

Math Nemotron 3 Ultra leads

Claude Opus 4.1: 22.3 (#277), Nemotron 3 Ultra: 35.3 (#187)

Math benchmarks
BenchmarkClaude Opus 4.1Nemotron 3 Ultra
OTIS Mock AIME 2024-202568.9%86.7%
LMArena Math14311457
FrontierMath (Tiers 1-3)12.6%—
FrontierMath Tier 42.4%—
ProofBench—2%
FrontierMath (Feb 2025 set)7.2%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Nemotron 3 Ultra leads

Claude Opus 4.1: 42.0 (#101), Nemotron 3 Ultra: 52.5 (#61)

Knowledge benchmarks
BenchmarkClaude Opus 4.1Nemotron 3 Ultra
GPQA Diamond77.3%85.4%
LMArena Expert14391472
Humanity's Last Exam11.5%—
Confabulations17.1%—
Vectara Hallucination Rate11.8%—

Multimodal Not comparable

Claude Opus 4.1: 26.8 (#119), Nemotron 3 Ultra: —

Multimodal benchmarks
BenchmarkClaude Opus 4.1Nemotron 3 Ultra
VPCT35%—

Multilingual Nemotron 3 Ultra leads

Claude Opus 4.1: 52.0 (#95), Nemotron 3 Ultra: 53.5 (#64)

Multilingual benchmarks
BenchmarkClaude Opus 4.1Nemotron 3 Ultra
LMArena Non-English14051427
LMArena Chinese14271497
LMArena French14311472
LMArena German14131471
LMArena Korean13801386
LMArena Russian14221417
LMArena Spanish14481454
LMArena Japanese1378—

Instruction Following Too close to call

Claude Opus 4.1: 75.6 (#58), Nemotron 3 Ultra: 74.7 (#88)

Instruction Following benchmarks
BenchmarkClaude Opus 4.1Nemotron 3 Ultra
LMArena Instruction Following14351418

Long Context Too close to call

Claude Opus 4.1: 44.5 (#63), Nemotron 3 Ultra: 43.9 (#81)

Long Context benchmarks
BenchmarkClaude Opus 4.1Nemotron 3 Ultra
LMArena Longer Query14551437

Writing & Preference Nemotron 3 Ultra leads

Claude Opus 4.1: 62.4 (#74), Nemotron 3 Ultra: 66.3 (#36)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.1Nemotron 3 Ultra
LMArena Text14191445
LMArena Creative Writing14121398
LMArena Multi-Turn14441411
Short-Story Creative Writing84.7%—
EQ-Bench Creative Writing—1692

Frequently asked questions

Is Claude Opus 4.1 better than Nemotron 3 Ultra?

Nemotron 3 Ultra is the stronger model overall, scoring 42.5 to 41.0 on the Noometry Index.

Which is cheaper, Claude Opus 4.1 or Nemotron 3 Ultra?

Nemotron 3 Ultra is cheaper. It lists at $0.50 per million input tokens and $2.20 per million output tokens; Claude Opus 4.1 lists at $15 and $75.

Is Claude Opus 4.1 or Nemotron 3 Ultra better for coding?

Claude Opus 4.1 scores higher on coding benchmarks: 44.4 versus 38.1 in the Noometry coding category.

Which has the bigger context window?

Nemotron 3 Ultra does, with 262K tokens against 200K.

How many benchmarks do Claude Opus 4.1 and Nemotron 3 Ultra share?

24 benchmarks have published results for both models. Claude Opus 4.1 has 48 scored results on Noometry and Nemotron 3 Ultra has 30.

Related comparisons

Go deeper