Model comparison
gpt-oss-20b vs Nova 2.0 Pro Preview
gpt-oss-20b and Nova 2.0 Pro Preview score almost the same on the Noometry Index (32.5 vs 33.4), so choose on price, context window or the category you care about most.
Last verified . 2 shared benchmarks.
Summary
- They share 2 benchmarks with published results for both. gpt-oss-20b scores higher in 0 categories and Nova 2.0 Pro Preview in 3 categories; 3 gaps are clear of the uncertainty.
- The widest gap is in agentic & tool use, where Nova 2.0 Pro Preview leads 21.9 to 9.3.
- The biggest single-benchmark swing is SciCode: 34.4% for gpt-oss-20b and 42.7% for Nova 2.0 Pro Preview.
- gpt-oss-20b has downloadable open weights; the other is API-only.
Side by side
| gpt-oss-20b | Nova 2.0 Pro Preview | |
|---|---|---|
| Provider | OpenAI | Amazon |
| Noometry Index | 32.5 | 33.4 |
| Released | 2025-08-05 | 2025-12-02 |
| Weights | Open | Proprietary |
| Context window | 131K | — |
| Max output | 16K | — |
| Input $ / M tokens | $0.018 | — |
| Output $ / M tokens | $0.09 | — |
| Results tracked | 34 | 3 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Nova 2.0 Pro Preview leads
gpt-oss-20b: 37.6 (#192), Nova 2.0 Pro Preview: 40.8 (#133)
| Benchmark | gpt-oss-20b | Nova 2.0 Pro Preview |
|---|---|---|
| SciCode | 34.4% | 42.7% |
| WeirdML | 40.9% | — |
| LMArena Coding | 1306 | — |
| ALE-Bench | 566.05 | — |
Agentic & Tool Use Nova 2.0 Pro Preview leads
gpt-oss-20b: 9.3 (#154), Nova 2.0 Pro Preview: 21.9 (#136)
| Benchmark | gpt-oss-20b | Nova 2.0 Pro Preview |
|---|---|---|
| Terminal-Bench | 3.4% | — |
| GDP.pdf | — | 2% |
Reasoning Nova 2.0 Pro Preview leads
gpt-oss-20b: 19.3 (#261), Nova 2.0 Pro Preview: 22.4 (#194)
| Benchmark | gpt-oss-20b | Nova 2.0 Pro Preview |
|---|---|---|
| CritPt | 1.4% | 0% |
| Kagi LLM Benchmark | 53.2% | — |
| Chess Puzzles | 4% | — |
| LMArena Hard Prompts | 1274 | — |
| DTBench | 68% | — |
| LMCA | 14.5% | — |
| Epoch Capabilities Index | 137.82 | — |
Math Not comparable
gpt-oss-20b: 39.4 (#103), Nova 2.0 Pro Preview: —
| Benchmark | gpt-oss-20b | Nova 2.0 Pro Preview |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 65.3% | — |
| Omni-MATH | 56.5% | — |
| LMArena Math | 1317 | — |
Knowledge Not comparable
gpt-oss-20b: 34.6 (#195), Nova 2.0 Pro Preview: —
| Benchmark | gpt-oss-20b | Nova 2.0 Pro Preview |
|---|---|---|
| GPQA Diamond | 60.8% | — |
| MMLU-Pro | 74% | — |
| GPQA (HELM) | 59.4% | — |
| LMArena Expert | 1258 | — |
Multilingual Not comparable
gpt-oss-20b: 42.2 (#197), Nova 2.0 Pro Preview: —
| Benchmark | gpt-oss-20b | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Non-English | 1268 | — |
| LMArena Chinese | 1314 | — |
| LMArena German | 1255 | — |
| LMArena Japanese | 1244 | — |
| LMArena Korean | 1236 | — |
| LMArena Russian | 1278 | — |
| LMArena Spanish | 1267 | — |
Instruction Following Not comparable
gpt-oss-20b: 61.8 (#240), Nova 2.0 Pro Preview: —
| Benchmark | gpt-oss-20b | Nova 2.0 Pro Preview |
|---|---|---|
| IFEval | 73.2% | — |
| LMArena Instruction Following | 1236 | — |
Long Context Not comparable
gpt-oss-20b: 37.9 (#209), Nova 2.0 Pro Preview: —
| Benchmark | gpt-oss-20b | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Longer Query | 1250 | — |
Writing & Preference Not comparable
gpt-oss-20b: 35.5 (#265), Nova 2.0 Pro Preview: —
| Benchmark | gpt-oss-20b | Nova 2.0 Pro Preview |
|---|---|---|
| LMArena Text | 1287 | — |
| LMArena Creative Writing | 1201 | — |
| EQ-Bench Creative Writing | 666 | — |
| WildBench | 73.7% | — |
| LMArena Multi-Turn | 1268 | — |
Frequently asked questions
Is gpt-oss-20b better than Nova 2.0 Pro Preview?
gpt-oss-20b and Nova 2.0 Pro Preview score almost the same on the Noometry Index (32.5 vs 33.4), so choose on price, context window or the category you care about most.
Is gpt-oss-20b or Nova 2.0 Pro Preview better for coding?
Nova 2.0 Pro Preview scores higher on coding benchmarks: 40.8 versus 37.6 in the Noometry coding category.
How many benchmarks do gpt-oss-20b and Nova 2.0 Pro Preview share?
2 benchmarks have published results for both models. gpt-oss-20b has 34 scored results on Noometry and Nova 2.0 Pro Preview has 3.