Model comparison

gpt-oss-20b vs Nova 2.0 Pro Preview

gpt-oss-20b and Nova 2.0 Pro Preview score almost the same on the Noometry Index (32.5 vs 33.4), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Nova 2.0 Pro Preview Amazon

33.4

Rank #244 Reported

Summary

  • They share 2 benchmarks with published results for both. gpt-oss-20b scores higher in 0 categories and Nova 2.0 Pro Preview in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Nova 2.0 Pro Preview leads 21.9 to 9.3.
  • The biggest single-benchmark swing is SciCode: 34.4% for gpt-oss-20b and 42.7% for Nova 2.0 Pro Preview.
  • gpt-oss-20b has downloadable open weights; the other is API-only.

Side by side

gpt-oss-20b and Nova 2.0 Pro Preview specifications
gpt-oss-20bNova 2.0 Pro Preview
ProviderOpenAIAmazon
Noometry Index32.533.4
Released2025-08-052025-12-02
WeightsOpenProprietary
Context window131K—
Max output16K—
Input $ / M tokens$0.018—
Output $ / M tokens$0.09—
Results tracked343

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nova 2.0 Pro Preview leads

gpt-oss-20b: 37.6 (#192), Nova 2.0 Pro Preview: 40.8 (#133)

Coding benchmarks
Benchmarkgpt-oss-20bNova 2.0 Pro Preview
SciCode34.4%42.7%
WeirdML40.9%—
LMArena Coding1306—
ALE-Bench566.05—

Agentic & Tool Use Nova 2.0 Pro Preview leads

gpt-oss-20b: 9.3 (#154), Nova 2.0 Pro Preview: 21.9 (#136)

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-20bNova 2.0 Pro Preview
Terminal-Bench3.4%—
GDP.pdf—2%

Reasoning Nova 2.0 Pro Preview leads

gpt-oss-20b: 19.3 (#261), Nova 2.0 Pro Preview: 22.4 (#194)

Reasoning benchmarks
Benchmarkgpt-oss-20bNova 2.0 Pro Preview
CritPt1.4%0%
Kagi LLM Benchmark53.2%—
Chess Puzzles4%—
LMArena Hard Prompts1274—
DTBench68%—
LMCA14.5%—
Epoch Capabilities Index137.82—

Math Not comparable

gpt-oss-20b: 39.4 (#103), Nova 2.0 Pro Preview: —

Math benchmarks
Benchmarkgpt-oss-20bNova 2.0 Pro Preview
OTIS Mock AIME 2024-202565.3%—
Omni-MATH56.5%—
LMArena Math1317—

Knowledge Not comparable

gpt-oss-20b: 34.6 (#195), Nova 2.0 Pro Preview: —

Knowledge benchmarks
Benchmarkgpt-oss-20bNova 2.0 Pro Preview
GPQA Diamond60.8%—
MMLU-Pro74%—
GPQA (HELM)59.4%—
LMArena Expert1258—

Multilingual Not comparable

gpt-oss-20b: 42.2 (#197), Nova 2.0 Pro Preview: —

Multilingual benchmarks
Benchmarkgpt-oss-20bNova 2.0 Pro Preview
LMArena Non-English1268—
LMArena Chinese1314—
LMArena German1255—
LMArena Japanese1244—
LMArena Korean1236—
LMArena Russian1278—
LMArena Spanish1267—

Instruction Following Not comparable

gpt-oss-20b: 61.8 (#240), Nova 2.0 Pro Preview: —

Instruction Following benchmarks
Benchmarkgpt-oss-20bNova 2.0 Pro Preview
IFEval73.2%—
LMArena Instruction Following1236—

Long Context Not comparable

gpt-oss-20b: 37.9 (#209), Nova 2.0 Pro Preview: —

Long Context benchmarks
Benchmarkgpt-oss-20bNova 2.0 Pro Preview
LMArena Longer Query1250—

Writing & Preference Not comparable

gpt-oss-20b: 35.5 (#265), Nova 2.0 Pro Preview: —

Writing & Preference benchmarks
Benchmarkgpt-oss-20bNova 2.0 Pro Preview
LMArena Text1287—
LMArena Creative Writing1201—
EQ-Bench Creative Writing666—
WildBench73.7%—
LMArena Multi-Turn1268—

Frequently asked questions

Is gpt-oss-20b better than Nova 2.0 Pro Preview?

gpt-oss-20b and Nova 2.0 Pro Preview score almost the same on the Noometry Index (32.5 vs 33.4), so choose on price, context window or the category you care about most.

Is gpt-oss-20b or Nova 2.0 Pro Preview better for coding?

Nova 2.0 Pro Preview scores higher on coding benchmarks: 40.8 versus 37.6 in the Noometry coding category.

How many benchmarks do gpt-oss-20b and Nova 2.0 Pro Preview share?

2 benchmarks have published results for both models. gpt-oss-20b has 34 scored results on Noometry and Nova 2.0 Pro Preview has 3.

Related comparisons

Go deeper