Model comparison

Nvidia Llama 3.3 Nemotron Super 49b v1.5 vs Step 3.7 Flash

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 37.3 on the Noometry Index.

Last verified . 0 shared benchmarks.

Step 3.7 Flash StepFun

37.3

Rank #207 Reported

Summary

  • The widest gap is in reasoning, where Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads 26.8 to 21.6.
  • Both cost about the same: $0.40 input and $0.40 output per million tokens.
  • Step 3.7 Flash accepts more context: 256K tokens versus 131K.

Side by side

Nvidia Llama 3.3 Nemotron Super 49b v1.5 and Step 3.7 Flash specifications
Nvidia Llama 3.3 Nemotron Super 49b v1.5Step 3.7 Flash
ProviderNVIDIAStepFun
Noometry Index40.337.3
Released2025-07-252026-05-29
WeightsOpenOpen
Context window131K256K
Max output131K256K
Input $ / M tokens$0.40$0.18
Output $ / M tokens$0.40$1.11
Results tracked125

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 39.8 (#154), Step 3.7 Flash: 40.0 (#150)

Coding benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Step 3.7 Flash
SciCode—40%
LMArena Coding1355—
ALE-Bench—694.12

Reasoning Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 26.8 (#128), Step 3.7 Flash: 21.6 (#219)

Reasoning benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Step 3.7 Flash
NYT Connections (extended)—39.7%
CritPt—2.3%
LMArena Hard Prompts1336—

Math Step 3.7 Flash leads

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 38.2 (#141), Step 3.7 Flash: 42.9 (#82)

Math benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Step 3.7 Flash
MathArena Final-Answer Competitions—68.5%
LMArena Math1392—

Knowledge Not comparable

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 36.7 (#165), Step 3.7 Flash: —

Knowledge benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Step 3.7 Flash
LMArena Expert1330—

Multilingual Not comparable

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 45.5 (#168), Step 3.7 Flash: —

Multilingual benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Step 3.7 Flash
LMArena Non-English1316—
LMArena Japanese1300—
LMArena Russian1332—

Instruction Following Not comparable

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 68.6 (#188), Step 3.7 Flash: —

Instruction Following benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Step 3.7 Flash
LMArena Instruction Following1299—

Long Context Not comparable

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 40.0 (#164), Step 3.7 Flash: —

Long Context benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Step 3.7 Flash
LMArena Longer Query1315—

Writing & Preference Not comparable

Nvidia Llama 3.3 Nemotron Super 49b v1.5: 53.1 (#159), Step 3.7 Flash: —

Writing & Preference benchmarks
BenchmarkNvidia Llama 3.3 Nemotron Super 49b v1.5Step 3.7 Flash
LMArena Text1338—
LMArena Creative Writing1307—
LMArena Multi-Turn1334—

Frequently asked questions

Is Nvidia Llama 3.3 Nemotron Super 49b v1.5 better than Step 3.7 Flash?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 37.3 on the Noometry Index.

Which is cheaper, Nvidia Llama 3.3 Nemotron Super 49b v1.5 or Step 3.7 Flash?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is cheaper. It lists at $0.40 per million input tokens and $0.40 per million output tokens; Step 3.7 Flash lists at $0.18 and $1.11.

Is Nvidia Llama 3.3 Nemotron Super 49b v1.5 or Step 3.7 Flash better for coding?

They score almost the same on coding (39.8 vs 40.0); test both on your own repository before choosing.

Which has the bigger context window?

Step 3.7 Flash does, with 256K tokens against 131K.

How many benchmarks do Nvidia Llama 3.3 Nemotron Super 49b v1.5 and Step 3.7 Flash share?

0 benchmarks have published results for both models. Nvidia Llama 3.3 Nemotron Super 49b v1.5 has 12 scored results on Noometry and Step 3.7 Flash has 5.

Related comparisons

Go deeper