Model comparison

Nemotron 3.5 Lightning vs o4-mini

o4-mini is the stronger model overall, scoring 41.6 to 40.0 on the Noometry Index. Nemotron 3.5 Lightning costs 22× less per token, which makes it the better buy when o4-mini's lead doesn't matter for your workload.

Last verified . 17 shared benchmarks.

Nemotron 3.5 Lightning NVIDIA

40.0

Rank #155 Confirmed

o4-mini OpenAI

41.6

Rank #132 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Nemotron 3.5 Lightning scores higher in 1 category and o4-mini in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o4-mini leads 43.6 to 37.5.
  • Nemotron 3.5 Lightning is cheaper at $0.05 / $0.20 per million input/output tokens, against $1.10 / $4.40 for o4-mini.
  • Nemotron 3.5 Lightning accepts more context: 262K tokens versus 200K.
  • Nemotron 3.5 Lightning has downloadable open weights; the other is API-only.

Side by side

Nemotron 3.5 Lightning and o4-mini specifications
Nemotron 3.5 Lightningo4-mini
ProviderNVIDIAOpenAI
Noometry Index40.041.6
Released2026-08-112025-04-16
WeightsOpenProprietary
Context window262K200K
Max output262K100K
Input $ / M tokens$0.05$1.10
Output $ / M tokens$0.20$4.40
Results tracked1860

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Nemotron 3.5 Lightning: 40.4 (#141), o4-mini: 40.9 (#127)

Coding benchmarks
BenchmarkNemotron 3.5 Lightningo4-mini
LMArena Coding13751368
SWE-bench Verified (bash only)—45%
Aider Polyglot—72%
GSO—3.6%
WeirdML—52.6%
CadEval—62%
ALE-Bench—826.17
AlgoTune—1.72

Agentic & Tool Use Not comparable

Nemotron 3.5 Lightning: —, o4-mini: 32.6 (#61)

Agentic & Tool Use benchmarks
BenchmarkNemotron 3.5 Lightningo4-mini
Berkeley Function Calling Leaderboard—53.2%
GDPval—25.3%
METR Time Horizons—63.9%

Reasoning Nemotron 3.5 Lightning leads

Nemotron 3.5 Lightning: 26.8 (#127), o4-mini: 24.6 (#162)

Reasoning benchmarks
BenchmarkNemotron 3.5 Lightningo4-mini
LMArena Hard Prompts13371351
ARC-AGI-2—6.1%
SimpleBench—38.7%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—58.7%
CritPt—0.6%
Chess Puzzles—26%
EnigmaEval—9.2%
Mystery Game Puzzles—5%
DTBench—77.6%
LMCA—26.5%
Epoch Capabilities Index—145.64
ForecastBench—61.8

Math o4-mini leads

Nemotron 3.5 Lightning: 37.5 (#155), o4-mini: 40.8 (#89)

Math benchmarks
BenchmarkNemotron 3.5 Lightningo4-mini
LMArena Math13591389
FrontierMath (Tiers 1-3)—36.1%
FrontierMath Tier 4—4.9%
OTIS Mock AIME 2024-2025—81.7%
Omni-MATH—72%
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—24.8%
FrontierMath Tier 4 (v1)—6.3%

Knowledge o4-mini leads

Nemotron 3.5 Lightning: 37.5 (#154), o4-mini: 43.6 (#91)

Knowledge benchmarks
BenchmarkNemotron 3.5 Lightningo4-mini
LMArena Expert13561343
GPQA Diamond—79.6%
Humanity's Last Exam—18.1%
SimpleQA Verified—19.6%
MMLU-Pro—82%
Confabulations—15.8%
Vectara Hallucination Rate—18.6%
GPQA (HELM)—73.5%

Multimodal Not comparable

Nemotron 3.5 Lightning: —, o4-mini: 40.2 (#49)

Multimodal benchmarks
BenchmarkNemotron 3.5 Lightningo4-mini
LMArena Vision—1194
GeoBench—64%
VPCT—57.5%

Multilingual o4-mini leads

Nemotron 3.5 Lightning: 44.0 (#180), o4-mini: 47.0 (#154)

Multilingual benchmarks
BenchmarkNemotron 3.5 Lightningo4-mini
LMArena Non-English12951337
LMArena Chinese13591354
LMArena French13661364
LMArena German12821336
LMArena Japanese12061308
LMArena Korean12381312
LMArena Russian12531334
LMArena Spanish13451347

Instruction Following o4-mini leads

Nemotron 3.5 Lightning: 69.6 (#170), o4-mini: 75.2 (#68)

Instruction Following benchmarks
BenchmarkNemotron 3.5 Lightningo4-mini
LMArena Instruction Following13181321
IFEval—92.8%

Long Context o4-mini leads

Nemotron 3.5 Lightning: 39.9 (#165), o4-mini: 45.5 (#33)

Long Context benchmarks
BenchmarkNemotron 3.5 Lightningo4-mini
LMArena Longer Query13141315
Fiction.LiveBench—77.8%

Writing & Preference o4-mini leads

Nemotron 3.5 Lightning: 48.5 (#201), o4-mini: 54.0 (#152)

Writing & Preference benchmarks
BenchmarkNemotron 3.5 Lightningo4-mini
LMArena Text13271353
LMArena Creative Writing12541294
LMArena Multi-Turn13281350
Short-Story Creative Writing—75%
EQ-Bench Creative Writing1280—
WildBench—85.4%

Frequently asked questions

Is Nemotron 3.5 Lightning better than o4-mini?

o4-mini is the stronger model overall, scoring 41.6 to 40.0 on the Noometry Index. Nemotron 3.5 Lightning costs 22× less per token, which makes it the better buy when o4-mini's lead doesn't matter for your workload.

Which is cheaper, Nemotron 3.5 Lightning or o4-mini?

Nemotron 3.5 Lightning is cheaper. It lists at $0.05 per million input tokens and $0.20 per million output tokens; o4-mini lists at $1.10 and $4.40.

Is Nemotron 3.5 Lightning or o4-mini better for coding?

They score almost the same on coding (40.4 vs 40.9); test both on your own repository before choosing.

Which has the bigger context window?

Nemotron 3.5 Lightning does, with 262K tokens against 200K.

How many benchmarks do Nemotron 3.5 Lightning and o4-mini share?

17 benchmarks have published results for both models. Nemotron 3.5 Lightning has 18 scored results on Noometry and o4-mini has 60.

Related comparisons

Go deeper