Model comparison

gpt-oss-120b vs Nemotron 3.5 Lightning

Nemotron 3.5 Lightning is the stronger model overall, scoring 40.0 to 36.3 on the Noometry Index.

Last verified . 18 shared benchmarks.

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Nemotron 3.5 Lightning NVIDIA

40.0

Rank #155 Confirmed

Summary

  • They share 18 benchmarks with published results for both. gpt-oss-120b scores higher in 3 categories and Nemotron 3.5 Lightning in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where gpt-oss-120b leads 52.5 to 37.5.
  • gpt-oss-120b is cheaper at $0.037 / $0.17 per million input/output tokens, against $0.05 / $0.20 for Nemotron 3.5 Lightning.
  • Nemotron 3.5 Lightning accepts more context: 262K tokens versus 131K.

Side by side

gpt-oss-120b and Nemotron 3.5 Lightning specifications
gpt-oss-120bNemotron 3.5 Lightning
ProviderOpenAINVIDIA
Noometry Index36.340.0
Released2025-08-052026-08-11
WeightsOpenOpen
Context window131K262K
Max output41K262K
Input $ / M tokens$0.037$0.05
Output $ / M tokens$0.17$0.20
Results tracked4818

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nemotron 3.5 Lightning leads

gpt-oss-120b: 33.5 (#256), Nemotron 3.5 Lightning: 40.4 (#141)

Coding benchmarks
Benchmarkgpt-oss-120bNemotron 3.5 Lightning
LMArena Coding13801375
SWE-bench Verified (bash only)26%—
Aider Polyglot41.8%—
SciCode36%—
WeirdML48.2%—
ALE-Bench575.62—
AlgoTune1.41—

Agentic & Tool Use Not comparable

gpt-oss-120b: 12.2 (#153), Nemotron 3.5 Lightning: —

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-120bNemotron 3.5 Lightning
Terminal-Bench18.7%—
APEX-Agents4.4%—
METR Time Horizons56.6%—
Vending-Bench 2-21.53—

Reasoning Nemotron 3.5 Lightning leads

gpt-oss-120b: 20.0 (#245), Nemotron 3.5 Lightning: 26.8 (#127)

Reasoning benchmarks
Benchmarkgpt-oss-120bNemotron 3.5 Lightning
LMArena Hard Prompts13641337
SimpleBench22.1%—
Kagi LLM Benchmark58.6%—
CritPt1.1%—
Chess Puzzles20%—
Mystery Game Puzzles2%—
DTBench76.3%—
LMCA22.1%—
Surface Evolver Bench25%—
Epoch Capabilities Index139.93—

Math gpt-oss-120b leads

gpt-oss-120b: 52.5 (#50), Nemotron 3.5 Lightning: 37.5 (#155)

Math benchmarks
Benchmarkgpt-oss-120bNemotron 3.5 Lightning
LMArena Math13891359
OTIS Mock AIME 2024-202588.9%—
Omni-MATH68.8%—

Knowledge gpt-oss-120b leads

gpt-oss-120b: 42.4 (#96), Nemotron 3.5 Lightning: 37.5 (#154)

Knowledge benchmarks
Benchmarkgpt-oss-120bNemotron 3.5 Lightning
LMArena Expert13561356
GPQA Diamond75.8%—
MMLU-Pro79.5%—
Confabulations15.7%—
Vectara Hallucination Rate14.2%—
GPQA (HELM)68.4%—

Multilingual gpt-oss-120b leads

gpt-oss-120b: 48.0 (#147), Nemotron 3.5 Lightning: 44.0 (#180)

Multilingual benchmarks
Benchmarkgpt-oss-120bNemotron 3.5 Lightning
LMArena Non-English13511295
LMArena Chinese13851359
LMArena French13691366
LMArena German13531282
LMArena Japanese13311206
LMArena Korean12821238
LMArena Russian13431253
LMArena Spanish13891345

Instruction Following Too close to call

gpt-oss-120b: 69.3 (#173), Nemotron 3.5 Lightning: 69.6 (#170)

Instruction Following benchmarks
Benchmarkgpt-oss-120bNemotron 3.5 Lightning
LMArena Instruction Following13181318
IFEval83.6%—

Long Context Nemotron 3.5 Lightning leads

gpt-oss-120b: 31.4 (#278), Nemotron 3.5 Lightning: 39.9 (#165)

Long Context benchmarks
Benchmarkgpt-oss-120bNemotron 3.5 Lightning
LMArena Longer Query13191314
Fiction.LiveBench44.4%—

Writing & Preference Nemotron 3.5 Lightning leads

gpt-oss-120b: 46.5 (#217), Nemotron 3.5 Lightning: 48.5 (#201)

Writing & Preference benchmarks
Benchmarkgpt-oss-120bNemotron 3.5 Lightning
LMArena Text13651327
LMArena Creative Writing12751254
EQ-Bench Creative Writing9611280
LMArena Multi-Turn13401328
Short-Story Creative Writing77.1%—
WildBench84.5%—

Frequently asked questions

Is gpt-oss-120b better than Nemotron 3.5 Lightning?

Nemotron 3.5 Lightning is the stronger model overall, scoring 40.0 to 36.3 on the Noometry Index.

Which is cheaper, gpt-oss-120b or Nemotron 3.5 Lightning?

gpt-oss-120b is cheaper. It lists at $0.037 per million input tokens and $0.17 per million output tokens; Nemotron 3.5 Lightning lists at $0.05 and $0.20.

Is gpt-oss-120b or Nemotron 3.5 Lightning better for coding?

Nemotron 3.5 Lightning scores higher on coding benchmarks: 40.4 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

Nemotron 3.5 Lightning does, with 262K tokens against 131K.

How many benchmarks do gpt-oss-120b and Nemotron 3.5 Lightning share?

18 benchmarks have published results for both models. gpt-oss-120b has 48 scored results on Noometry and Nemotron 3.5 Lightning has 18.

Related comparisons

Go deeper