Model comparison

DeepSeek-V3 vs gpt-oss-20b

DeepSeek-V3 is the stronger model overall, scoring 39.5 to 32.5 on the Noometry Index. gpt-oss-20b costs 11× less per token, which makes it the better buy when DeepSeek-V3's lead doesn't matter for your workload.

Last verified . 31 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Summary

  • They share 31 benchmarks with published results for both. DeepSeek-V3 scores higher in 6 categories and gpt-oss-20b in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3 leads 57.4 to 35.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 37.8% for DeepSeek-V3 and 65.3% for gpt-oss-20b.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $0.24 / $0.90 for DeepSeek-V3.
  • DeepSeek-V3 accepts more context: 164K tokens versus 131K.

Side by side

DeepSeek-V3 and gpt-oss-20b specifications
DeepSeek-V3gpt-oss-20b
ProviderDeepSeekOpenAI
Noometry Index39.532.5
Released2024-12-262025-08-05
WeightsOpenOpen
Context window164K131K
Max output164K16K
Input $ / M tokens$0.24$0.018
Output $ / M tokens$0.90$0.09
Results tracked6034

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), gpt-oss-20b: 37.6 (#192)

Coding benchmarks
BenchmarkDeepSeek-V3gpt-oss-20b
SciCode35.8%34.4%
WeirdML36.1%40.9%
LMArena Coding13681306
Aider Polyglot55.1%—
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
ALE-Bench—566.05
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, gpt-oss-20b: 9.3 (#154)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3gpt-oss-20b
Terminal-Bench—3.4%
METR Time Horizons49.6%—

Reasoning DeepSeek-V3 leads

DeepSeek-V3: 20.5 (#236), gpt-oss-20b: 19.3 (#261)

Reasoning benchmarks
BenchmarkDeepSeek-V3gpt-oss-20b
Kagi LLM Benchmark52.3%53.2%
CritPt0%1.4%
LMArena Hard Prompts13651274
DTBench64.8%68%
LMCA15.5%14.5%
Epoch Capabilities Index135.94137.82
SimpleBench27.2%—
Chess Puzzles—4%
LiveBench Reasoning65.8%—
LiveBench Data Analysis60.9%—
BIG-Bench Hard87.5%—
ForecastBench59.1—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math gpt-oss-20b leads

DeepSeek-V3: 32.1 (#219), gpt-oss-20b: 39.4 (#103)

Math benchmarks
BenchmarkDeepSeek-V3gpt-oss-20b
OTIS Mock AIME 2024-202537.8%65.3%
Omni-MATH40.3%56.5%
LMArena Math13731317
LiveBench Math73.5%—
MATH Level 575.5%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge DeepSeek-V3 leads

DeepSeek-V3: 37.5 (#155), gpt-oss-20b: 34.6 (#195)

Knowledge benchmarks
BenchmarkDeepSeek-V3gpt-oss-20b
GPQA Diamond67.6%60.8%
MMLU-Pro72.3%74%
GPQA (HELM)53.8%59.4%
LMArena Expert13511258
Confabulations26.1%—
Vectara Hallucination Rate6.1%—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multilingual DeepSeek-V3 leads

DeepSeek-V3: 48.5 (#143), gpt-oss-20b: 42.2 (#197)

Multilingual benchmarks
BenchmarkDeepSeek-V3gpt-oss-20b
LMArena Non-English13581268
LMArena Chinese13911314
LMArena German13741255
LMArena Japanese13331244
LMArena Korean13191236
LMArena Russian13731278
LMArena Spanish13581267
LMArena French1385—

Instruction Following DeepSeek-V3 leads

DeepSeek-V3: 72.8 (#130), gpt-oss-20b: 61.8 (#240)

Instruction Following benchmarks
BenchmarkDeepSeek-V3gpt-oss-20b
IFEval83.2%73.2%
LMArena Instruction Following13451236
LiveBench Instruction Following81.5%—

Long Context gpt-oss-20b leads

DeepSeek-V3: 34.0 (#253), gpt-oss-20b: 37.9 (#209)

Long Context benchmarks
BenchmarkDeepSeek-V3gpt-oss-20b
LMArena Longer Query13521250
Fiction.LiveBench50%—

Writing & Preference DeepSeek-V3 leads

DeepSeek-V3: 57.4 (#130), gpt-oss-20b: 35.5 (#265)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3gpt-oss-20b
LMArena Text13751287
LMArena Creative Writing13641201
EQ-Bench Creative Writing1472666
WildBench83%73.7%
LMArena Multi-Turn13891268
Short-Story Creative Writing77%—
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than gpt-oss-20b?

DeepSeek-V3 is the stronger model overall, scoring 39.5 to 32.5 on the Noometry Index. gpt-oss-20b costs 11× less per token, which makes it the better buy when DeepSeek-V3's lead doesn't matter for your workload.

Which is cheaper, DeepSeek-V3 or gpt-oss-20b?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; DeepSeek-V3 lists at $0.24 and $0.90.

Is DeepSeek-V3 or gpt-oss-20b better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 37.6 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3 does, with 164K tokens against 131K.

How many benchmarks do DeepSeek-V3 and gpt-oss-20b share?

31 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and gpt-oss-20b has 34.

Related comparisons

Go deeper