Model comparison

Seed 2.0 Pro vs gpt-oss-120b

Seed 2.0 Pro is the stronger model overall, scoring 43.2 to 36.3 on the Noometry Index. gpt-oss-120b costs 16× less per token, which makes it the better buy when Seed 2.0 Pro's lead doesn't matter for your workload.

Last verified . 17 shared benchmarks.

Seed 2.0 Pro ByteDance Seed

43.2

Rank #96 Confirmed

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Seed 2.0 Pro scores higher in 6 categories and gpt-oss-120b in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Seed 2.0 Pro leads 62.9 to 46.5.
  • gpt-oss-120b is cheaper at $0.037 / $0.17 per million input/output tokens, against $0.50 / $3 for Seed 2.0 Pro.
  • Seed 2.0 Pro accepts more context: 256K tokens versus 131K.
  • gpt-oss-120b has downloadable open weights; the other is API-only.

Side by side

Seed 2.0 Pro and gpt-oss-120b specifications
Seed 2.0 Progpt-oss-120b
ProviderByteDance SeedOpenAI
Noometry Index43.236.3
Released2026-02-142025-08-05
WeightsProprietaryOpen
Context window256K131K
Max output128K41K
Input $ / M tokens$0.50$0.037
Output $ / M tokens$3$0.17
Results tracked2048

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Seed 2.0 Pro leads

Seed 2.0 Pro: 43.5 (#86), gpt-oss-120b: 33.5 (#256)

Coding benchmarks
BenchmarkSeed 2.0 Progpt-oss-120b
LMArena Coding14721380
SWE-bench Verified (bash only)—26%
Aider Polyglot—41.8%
SciCode—36%
WeirdML—48.2%
ALE-Bench—575.62
AlgoTune—1.41

Agentic & Tool Use Not comparable

Seed 2.0 Pro: —, gpt-oss-120b: 12.2 (#153)

Agentic & Tool Use benchmarks
BenchmarkSeed 2.0 Progpt-oss-120b
Terminal-Bench—18.7%
APEX-Agents—4.4%
METR Time Horizons—56.6%
Vending-Bench 2—-21.53

Reasoning Seed 2.0 Pro leads

Seed 2.0 Pro: 24.1 (#165), gpt-oss-120b: 20.0 (#245)

Reasoning benchmarks
BenchmarkSeed 2.0 Progpt-oss-120b
LMArena Hard Prompts14531364
SimpleBench—22.1%
Kagi LLM Benchmark—58.6%
NYT Connections (extended)28.4%—
CritPt—1.1%
Chess Puzzles—20%
Thematic Generalization57.1%—
Mystery Game Puzzles—2%
DTBench—76.3%
LMCA—22.1%
Surface Evolver Bench—25%
Epoch Capabilities Index—139.93

Math gpt-oss-120b leads

Seed 2.0 Pro: 39.3 (#108), gpt-oss-120b: 52.5 (#50)

Math benchmarks
BenchmarkSeed 2.0 Progpt-oss-120b
LMArena Math14391389
OTIS Mock AIME 2024-2025—88.9%
Omni-MATH—68.8%

Knowledge gpt-oss-120b leads

Seed 2.0 Pro: 40.2 (#122), gpt-oss-120b: 42.4 (#96)

Knowledge benchmarks
BenchmarkSeed 2.0 Progpt-oss-120b
LMArena Expert14401356
GPQA Diamond—75.8%
MMLU-Pro—79.5%
Confabulations—15.7%
Vectara Hallucination Rate—14.2%
GPQA (HELM)—68.4%

Multimodal Not comparable

Seed 2.0 Pro: 41.5 (#35), gpt-oss-120b: —

Multimodal benchmarks
BenchmarkSeed 2.0 Progpt-oss-120b
LMArena Vision1274—

Multilingual Seed 2.0 Pro leads

Seed 2.0 Pro: 54.5 (#39), gpt-oss-120b: 48.0 (#147)

Multilingual benchmarks
BenchmarkSeed 2.0 Progpt-oss-120b
LMArena Non-English14411351
LMArena Chinese14891385
LMArena French14711369
LMArena German14421353
LMArena Japanese14081331
LMArena Korean14111282
LMArena Russian14491343
LMArena Spanish14601389

Instruction Following Seed 2.0 Pro leads

Seed 2.0 Pro: 74.5 (#91), gpt-oss-120b: 69.3 (#173)

Instruction Following benchmarks
BenchmarkSeed 2.0 Progpt-oss-120b
LMArena Instruction Following14141318
IFEval—83.6%

Long Context Seed 2.0 Pro leads

Seed 2.0 Pro: 43.6 (#90), gpt-oss-120b: 31.4 (#278)

Long Context benchmarks
BenchmarkSeed 2.0 Progpt-oss-120b
LMArena Longer Query14281319
Fiction.LiveBench—44.4%

Writing & Preference Seed 2.0 Pro leads

Seed 2.0 Pro: 62.9 (#69), gpt-oss-120b: 46.5 (#217)

Writing & Preference benchmarks
BenchmarkSeed 2.0 Progpt-oss-120b
LMArena Text14481365
LMArena Creative Writing14061275
LMArena Multi-Turn14411340
Short-Story Creative Writing—77.1%
EQ-Bench Creative Writing—961
WildBench—84.5%

Frequently asked questions

Is Seed 2.0 Pro better than gpt-oss-120b?

Seed 2.0 Pro is the stronger model overall, scoring 43.2 to 36.3 on the Noometry Index. gpt-oss-120b costs 16× less per token, which makes it the better buy when Seed 2.0 Pro's lead doesn't matter for your workload.

Which is cheaper, Seed 2.0 Pro or gpt-oss-120b?

gpt-oss-120b is cheaper. It lists at $0.037 per million input tokens and $0.17 per million output tokens; Seed 2.0 Pro lists at $0.50 and $3.

Is Seed 2.0 Pro or gpt-oss-120b better for coding?

Seed 2.0 Pro scores higher on coding benchmarks: 43.5 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

Seed 2.0 Pro does, with 256K tokens against 131K.

How many benchmarks do Seed 2.0 Pro and gpt-oss-120b share?

17 benchmarks have published results for both models. Seed 2.0 Pro has 20 scored results on Noometry and gpt-oss-120b has 48.

Related comparisons

Go deeper