Model comparison

Nvidia Llama 3.3 Nemotron Super 49b v1.5 vs Olmo 3.1 32b Think

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 37.9 on the Noometry Index.

Last verified . 11 shared benchmarks.

Summary

  • They share 11 benchmarks with published results for both. Nvidia Llama 3.3 Nemotron Super 49b v1.5 scores higher in 8 categories and Olmo 3.1 32b Think in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads 45.5 to 38.1.

Side by side

Nvidia Llama 3.3 Nemotron Super 49b v1.5 and Olmo 3.1 32b Think specifications
Nvidia Llama 3.3 Nemotron Super 49b v1.5Olmo 3.1 32b Think
ProviderNVIDIAAllen Institute for AI (Ai2)
Noometry Index40.337.9
Released2025-07-25—
WeightsOpenOpen
Context window131K—
Max output131K—
Input $ / M tokens$0.40—
Output $ / M tokens$0.40—
Results tracked1215

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 39.8 (#154), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Olmo 3.1 32b Think
LMArena Coding13551291

Reasoning Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 26.8 (#128), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Olmo 3.1 32b Think
LMArena Hard Prompts13361272

Math Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 38.2 (#141), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Olmo 3.1 32b Think
LMArena Math13921305

Knowledge Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 36.7 (#165), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Olmo 3.1 32b Think
LMArena Expert13301295

Multilingual Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 45.5 (#168), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Olmo 3.1 32b Think
LMArena Non-English13161209
LMArena Russian13321193
LMArena Chinese—1242
LMArena French—1260
LMArena German—1262
LMArena Japanese1300—
LMArena Spanish—1289

Instruction Following Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 68.6 (#188), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Olmo 3.1 32b Think
LMArena Instruction Following12991247

Long Context Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 40.0 (#164), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Olmo 3.1 32b Think
LMArena Longer Query13151272

Writing & Preference Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 53.1 (#159), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Olmo 3.1 32b Think
LMArena Text13381272
LMArena Creative Writing13071226
LMArena Multi-Turn13341252

Frequently asked questions

Is Nvidia Llama 3.3 Nemotron Super 49b v1.5 better than Olmo 3.1 32b Think?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 37.9 on the Noometry Index.

Is Nvidia Llama 3.3 Nemotron Super 49b v1.5 or Olmo 3.1 32b Think better for coding?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 scores higher on coding benchmarks: 39.8 versus 37.7 in the Noometry coding category.

How many benchmarks do Nvidia Llama 3.3 Nemotron Super 49b v1.5 and Olmo 3.1 32b Think share?

11 benchmarks have published results for both models. Nvidia Llama 3.3 Nemotron Super 49b v1.5 has 12 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper