Model comparison

Claude Haiku 4.5 vs Nvidia Llama 3.3 Nemotron Super 49b v1.5

Claude Haiku 4.5 and Nvidia Llama 3.3 Nemotron Super 49b v1.5 score almost the same on the Noometry Index (39.5 vs 40.3), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

Claude Haiku 4.5 Anthropic

39.5

Rank #165 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Claude Haiku 4.5 scores higher in 7 categories and Nvidia Llama 3.3 Nemotron Super 49b v1.5 in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads 26.8 to 15.1.
  • Nvidia Llama 3.3 Nemotron Super 49b v1.5 is cheaper at $0.40 / $0.40 per million input/output tokens, against $1 / $5 for Claude Haiku 4.5.
  • Claude Haiku 4.5 accepts more context: 200K tokens versus 131K.
  • Nvidia Llama 3.3 Nemotron Super 49b v1.5 has downloadable open weights; the other is API-only.

Side by side

Claude Haiku 4.5 and Nvidia Llama 3.3 Nemotron Super 49b v1.5 specifications
Claude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
ProviderAnthropicNVIDIA
Noometry Index39.540.3
Released2025-10-152025-07-25
WeightsProprietaryOpen
Context window200K131K
Max output64K131K
Input $ / M tokens$1$0.40
Output $ / M tokens$5$0.40
Results tracked5312

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Haiku 4.5 leads

Claude Haiku 4.5: 44.0 (#78), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 39.8 (#154)

Coding benchmarks
BenchmarkClaude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Coding14531355
SWE-bench Verified (bash only)66.6%—
LMArena WebDev1330—
SWE-bench Multilingual64.7%—
SciCode43.3%—
WeirdML45.4%—
ALE-Bench653.48—

Agentic & Tool Use Not comparable

Claude Haiku 4.5: 33.6 (#52), Nvidia Llama 3.3 Nemotron Super 49b v1.5: —

Agentic & Tool Use benchmarks
BenchmarkClaude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
Terminal-Bench35.5%—
Berkeley Function Calling Leaderboard68.7%—
DeepResearch Bench45.5%—
BALROG31.2%—
ExploitBench13.7%—
Vending-Bench 2458.89—

Reasoning Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Claude Haiku 4.5: 15.1 (#320), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 26.8 (#128)

Reasoning benchmarks
BenchmarkClaude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Hard Prompts14201336
ARC-AGI-24%—
NYT Connections (extended)14.3%—
ARC-AGI-147.7%—
CritPt0%—
Chess Puzzles8%—
DTBench73.6%—
LMCA30.9%—
Epoch Capabilities Index142.41—
ForecastBench61.4—

Math Claude Haiku 4.5 leads

Claude Haiku 4.5: 44.9 (#78), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 38.2 (#141)

Math benchmarks
BenchmarkClaude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Math13961392
OTIS Mock AIME 2024-202566.7%—
Omni-MATH56.1%—
MATH Level 596.4%—
FrontierMath (Feb 2025 set)5.9%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge Too close to call

Claude Haiku 4.5: 37.7 (#153), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 36.7 (#165)

Knowledge benchmarks
BenchmarkClaude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Expert14421330
GPQA Diamond71.2%—
SimpleQA Verified13.2%—
MMLU-Pro77.7%—
Vectara Hallucination Rate9.8%—
GPQA (HELM)60.5%—

Multimodal Not comparable

Claude Haiku 4.5: 26.8 (#118), Nvidia Llama 3.3 Nemotron Super 49b v1.5: —

Multimodal benchmarks
BenchmarkClaude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
Blueprint-Bench 20%—
LMArena Document1420—

Multilingual Claude Haiku 4.5 leads

Claude Haiku 4.5: 49.9 (#129), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 45.5 (#168)

Multilingual benchmarks
BenchmarkClaude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Non-English13771316
LMArena Japanese13391300
LMArena Russian13811332
LMArena Chinese1417—
LMArena French1408—
LMArena German1375—
LMArena Korean1347—
LMArena Spanish1420—

Instruction Following Claude Haiku 4.5 leads

Claude Haiku 4.5: 71.4 (#149), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 68.6 (#188)

Instruction Following benchmarks
BenchmarkClaude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Instruction Following14141299
IFEval80.1%—

Long Context Claude Haiku 4.5 leads

Claude Haiku 4.5: 43.6 (#92), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 40.0 (#164)

Long Context benchmarks
BenchmarkClaude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Longer Query14271315

Writing & Preference Claude Haiku 4.5 leads

Claude Haiku 4.5: 57.9 (#123), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 53.1 (#159)

Writing & Preference benchmarks
BenchmarkClaude Haiku 4.5Nvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Text13961338
LMArena Creative Writing13721307
LMArena Multi-Turn14091334
WildBench83.9%—
EQ-Bench 41064—

Frequently asked questions

Is Claude Haiku 4.5 better than Nvidia Llama 3.3 Nemotron Super 49b v1.5?

Claude Haiku 4.5 and Nvidia Llama 3.3 Nemotron Super 49b v1.5 score almost the same on the Noometry Index (39.5 vs 40.3), so choose on price, context window or the category you care about most.

Which is cheaper, Claude Haiku 4.5 or Nvidia Llama 3.3 Nemotron Super 49b v1.5?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is cheaper. It lists at $0.40 per million input tokens and $0.40 per million output tokens; Claude Haiku 4.5 lists at $1 and $5.

Is Claude Haiku 4.5 or Nvidia Llama 3.3 Nemotron Super 49b v1.5 better for coding?

Claude Haiku 4.5 scores higher on coding benchmarks: 44.0 versus 39.8 in the Noometry coding category.

Which has the bigger context window?

Claude Haiku 4.5 does, with 200K tokens against 131K.

How many benchmarks do Claude Haiku 4.5 and Nvidia Llama 3.3 Nemotron Super 49b v1.5 share?

12 benchmarks have published results for both models. Claude Haiku 4.5 has 53 scored results on Noometry and Nvidia Llama 3.3 Nemotron Super 49b v1.5 has 12.

Related comparisons

Go deeper