Model comparison

Seed 2.0 Pro vs gpt-oss-20b

Seed 2.0 Pro is the stronger model overall, scoring 43.2 to 32.5 on the Noometry Index. gpt-oss-20b costs 31× less per token, which makes it the better buy when Seed 2.0 Pro's lead doesn't matter for your workload.

Last verified . 16 shared benchmarks.

Seed 2.0 Pro ByteDance Seed

43.2

Rank #96 Confirmed

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Seed 2.0 Pro scores higher in 7 categories and gpt-oss-20b in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Seed 2.0 Pro leads 62.9 to 35.5.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $0.50 / $3 for Seed 2.0 Pro.
  • Seed 2.0 Pro accepts more context: 256K tokens versus 131K.
  • gpt-oss-20b has downloadable open weights; the other is API-only.

Side by side

Seed 2.0 Pro and gpt-oss-20b specifications
Seed 2.0 Progpt-oss-20b
ProviderByteDance SeedOpenAI
Noometry Index43.232.5
Released2026-02-142025-08-05
WeightsProprietaryOpen
Context window256K131K
Max output128K16K
Input $ / M tokens$0.50$0.018
Output $ / M tokens$3$0.09
Results tracked2034

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Seed 2.0 Pro leads

Seed 2.0 Pro: 43.5 (#86), gpt-oss-20b: 37.6 (#192)

Coding benchmarks
BenchmarkSeed 2.0 Progpt-oss-20b
LMArena Coding14721306
SciCode—34.4%
WeirdML—40.9%
ALE-Bench—566.05

Agentic & Tool Use Not comparable

Seed 2.0 Pro: —, gpt-oss-20b: 9.3 (#154)

Agentic & Tool Use benchmarks
BenchmarkSeed 2.0 Progpt-oss-20b
Terminal-Bench—3.4%

Reasoning Seed 2.0 Pro leads

Seed 2.0 Pro: 24.1 (#165), gpt-oss-20b: 19.3 (#261)

Reasoning benchmarks
BenchmarkSeed 2.0 Progpt-oss-20b
LMArena Hard Prompts14531274
Kagi LLM Benchmark—53.2%
NYT Connections (extended)28.4%—
CritPt—1.4%
Chess Puzzles—4%
Thematic Generalization57.1%—
DTBench—68%
LMCA—14.5%
Epoch Capabilities Index—137.82

Math Too close to call

Seed 2.0 Pro: 39.3 (#108), gpt-oss-20b: 39.4 (#103)

Math benchmarks
BenchmarkSeed 2.0 Progpt-oss-20b
LMArena Math14391317
OTIS Mock AIME 2024-2025—65.3%
Omni-MATH—56.5%

Knowledge Seed 2.0 Pro leads

Seed 2.0 Pro: 40.2 (#122), gpt-oss-20b: 34.6 (#195)

Knowledge benchmarks
BenchmarkSeed 2.0 Progpt-oss-20b
LMArena Expert14401258
GPQA Diamond—60.8%
MMLU-Pro—74%
GPQA (HELM)—59.4%

Multimodal Not comparable

Seed 2.0 Pro: 41.5 (#35), gpt-oss-20b: —

Multimodal benchmarks
BenchmarkSeed 2.0 Progpt-oss-20b
LMArena Vision1274—

Multilingual Seed 2.0 Pro leads

Seed 2.0 Pro: 54.5 (#39), gpt-oss-20b: 42.2 (#197)

Multilingual benchmarks
BenchmarkSeed 2.0 Progpt-oss-20b
LMArena Non-English14411268
LMArena Chinese14891314
LMArena German14421255
LMArena Japanese14081244
LMArena Korean14111236
LMArena Russian14491278
LMArena Spanish14601267
LMArena French1471—

Instruction Following Seed 2.0 Pro leads

Seed 2.0 Pro: 74.5 (#91), gpt-oss-20b: 61.8 (#240)

Instruction Following benchmarks
BenchmarkSeed 2.0 Progpt-oss-20b
LMArena Instruction Following14141236
IFEval—73.2%

Long Context Seed 2.0 Pro leads

Seed 2.0 Pro: 43.6 (#90), gpt-oss-20b: 37.9 (#209)

Long Context benchmarks
BenchmarkSeed 2.0 Progpt-oss-20b
LMArena Longer Query14281250

Writing & Preference Seed 2.0 Pro leads

Seed 2.0 Pro: 62.9 (#69), gpt-oss-20b: 35.5 (#265)

Writing & Preference benchmarks
BenchmarkSeed 2.0 Progpt-oss-20b
LMArena Text14481287
LMArena Creative Writing14061201
LMArena Multi-Turn14411268
EQ-Bench Creative Writing—666
WildBench—73.7%

Frequently asked questions

Is Seed 2.0 Pro better than gpt-oss-20b?

Seed 2.0 Pro is the stronger model overall, scoring 43.2 to 32.5 on the Noometry Index. gpt-oss-20b costs 31× less per token, which makes it the better buy when Seed 2.0 Pro's lead doesn't matter for your workload.

Which is cheaper, Seed 2.0 Pro or gpt-oss-20b?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Seed 2.0 Pro lists at $0.50 and $3.

Is Seed 2.0 Pro or gpt-oss-20b better for coding?

Seed 2.0 Pro scores higher on coding benchmarks: 43.5 versus 37.6 in the Noometry coding category.

Which has the bigger context window?

Seed 2.0 Pro does, with 256K tokens against 131K.

How many benchmarks do Seed 2.0 Pro and gpt-oss-20b share?

16 benchmarks have published results for both models. Seed 2.0 Pro has 20 scored results on Noometry and gpt-oss-20b has 34.

Related comparisons

Go deeper