Model comparison

gpt-oss-20b vs Nvidia Llama 3.3 Nemotron Super 49b v1.5

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 32.5 on the Noometry Index. gpt-oss-20b costs 11× less per token, which makes it the better buy when Nvidia Llama 3.3 Nemotron Super 49b v1.5's lead doesn't matter for your workload.

Last verified . 12 shared benchmarks.

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Summary

  • They share 12 benchmarks with published results for both. gpt-oss-20b scores higher in 1 category and Nvidia Llama 3.3 Nemotron Super 49b v1.5 in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads 53.1 to 35.5.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $0.40 / $0.40 for Nvidia Llama 3.3 Nemotron Super 49b v1.5.

Side by side

gpt-oss-20b and Nvidia Llama 3.3 Nemotron Super 49b v1.5 specifications
gpt-oss-20bNvidia Llama 3.3 Nemotron Super 49b v1.5
ProviderOpenAINVIDIA
Noometry Index32.540.3
Released2025-08-052025-07-25
WeightsOpenOpen
Context window131K131K
Max output16K131K
Input $ / M tokens$0.018$0.40
Output $ / M tokens$0.09$0.40
Results tracked3412

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

gpt-oss-20b: 37.6 (#192), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 39.8 (#154)

Coding benchmarks
Benchmarkgpt-oss-20bNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Coding13061355
SciCode34.4%—
WeirdML40.9%—
ALE-Bench566.05—

Agentic & Tool Use Not comparable

gpt-oss-20b: 9.3 (#154), Nvidia Llama 3.3 Nemotron Super 49b v1.5: —

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-20bNvidia Llama 3.3 Nemotron Super 49b v1.5
Terminal-Bench3.4%—

Reasoning Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

gpt-oss-20b: 19.3 (#261), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 26.8 (#128)

Reasoning benchmarks
Benchmarkgpt-oss-20bNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Hard Prompts12741336
Kagi LLM Benchmark53.2%—
CritPt1.4%—
Chess Puzzles4%—
DTBench68%—
LMCA14.5%—
Epoch Capabilities Index137.82—

Math gpt-oss-20b leads

gpt-oss-20b: 39.4 (#103), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 38.2 (#141)

Math benchmarks
Benchmarkgpt-oss-20bNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Math13171392
OTIS Mock AIME 2024-202565.3%—
Omni-MATH56.5%—

Knowledge Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

gpt-oss-20b: 34.6 (#195), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 36.7 (#165)

Knowledge benchmarks
Benchmarkgpt-oss-20bNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Expert12581330
GPQA Diamond60.8%—
MMLU-Pro74%—
GPQA (HELM)59.4%—

Multilingual Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

gpt-oss-20b: 42.2 (#197), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 45.5 (#168)

Multilingual benchmarks
Benchmarkgpt-oss-20bNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Non-English12681316
LMArena Japanese12441300
LMArena Russian12781332
LMArena Chinese1314—
LMArena German1255—
LMArena Korean1236—
LMArena Spanish1267—

Instruction Following Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

gpt-oss-20b: 61.8 (#240), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 68.6 (#188)

Instruction Following benchmarks
Benchmarkgpt-oss-20bNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Instruction Following12361299
IFEval73.2%—

Long Context Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

gpt-oss-20b: 37.9 (#209), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 40.0 (#164)

Long Context benchmarks
Benchmarkgpt-oss-20bNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Longer Query12501315

Writing & Preference Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

gpt-oss-20b: 35.5 (#265), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 53.1 (#159)

Writing & Preference benchmarks
Benchmarkgpt-oss-20bNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Text12871338
LMArena Creative Writing12011307
LMArena Multi-Turn12681334
EQ-Bench Creative Writing666—
WildBench73.7%—

Frequently asked questions

Is gpt-oss-20b better than Nvidia Llama 3.3 Nemotron Super 49b v1.5?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 32.5 on the Noometry Index. gpt-oss-20b costs 11× less per token, which makes it the better buy when Nvidia Llama 3.3 Nemotron Super 49b v1.5's lead doesn't matter for your workload.

Which is cheaper, gpt-oss-20b or Nvidia Llama 3.3 Nemotron Super 49b v1.5?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Nvidia Llama 3.3 Nemotron Super 49b v1.5 lists at $0.40 and $0.40.

Is gpt-oss-20b or Nvidia Llama 3.3 Nemotron Super 49b v1.5 better for coding?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 scores higher on coding benchmarks: 39.8 versus 37.6 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do gpt-oss-20b and Nvidia Llama 3.3 Nemotron Super 49b v1.5 share?

12 benchmarks have published results for both models. gpt-oss-20b has 34 scored results on Noometry and Nvidia Llama 3.3 Nemotron Super 49b v1.5 has 12.

Related comparisons

Go deeper