Model comparison

DeepSeek LLM 67B vs o3

o3 is the stronger model overall, scoring 47.5 to 24.9 on the Noometry Index.

Last verified . 15 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 15 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 0 categories and o3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 7.0.
  • The biggest single-benchmark swing is MATH Level 5: 6.4% for DeepSeek LLM 67B and 97.8% for o3.
  • DeepSeek LLM 67B has downloadable open weights; the other is API-only.

Side by side

DeepSeek LLM 67B and o3 specifications
DeepSeek LLM 67Bo3
ProviderDeepSeekOpenAI
Noometry Index24.947.5
Released2023-11-292025-04-16
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked1563

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

DeepSeek LLM 67B: 31.9 (#278), o3: 46.8 (#64)

Coding benchmarks
BenchmarkDeepSeek LLM 67Bo3
LMArena Coding10961408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

DeepSeek LLM 67B: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek LLM 67Bo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

DeepSeek LLM 67B: 16.5 (#304), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67Bo3
Chess Puzzles0%38%
LMArena Hard Prompts10701402
Epoch Capabilities Index110.5146.86
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
ForecastBench—62.5

Math o3 leads

DeepSeek LLM 67B: 8.7 (#324), o3: 50.2 (#58)

Math benchmarks
BenchmarkDeepSeek LLM 67Bo3
OTIS Mock AIME 2024-20250.8%84.4%
LMArena Math11081426
MATH Level 56.4%97.8%
FrontierMath (Tiers 1-3)—33.3%
Omni-MATH—71.4%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

DeepSeek LLM 67B: 7.0 (#313), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67Bo3
GPQA Diamond24.6%81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%
LMArena Expert—1402

Multimodal Not comparable

DeepSeek LLM 67B: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkDeepSeek LLM 67Bo3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual o3 leads

DeepSeek LLM 67B: 29.4 (#267), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67Bo3
LMArena Non-English10731401
LMArena Chinese11321437
LMArena French—1430
LMArena German—1420
LMArena Japanese—1403
LMArena Korean—1370
LMArena Russian—1406
LMArena Spanish—1395

Instruction Following o3 leads

DeepSeek LLM 67B: 55.4 (#277), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67Bo3
LMArena Instruction Following10791368
IFEval—86.9%

Long Context o3 leads

DeepSeek LLM 67B: 33.1 (#265), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkDeepSeek LLM 67Bo3
LMArena Longer Query10921372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

DeepSeek LLM 67B: 31.6 (#282), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67Bo3
LMArena Text11051410
LMArena Creative Writing10671359
LMArena Multi-Turn10821405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is DeepSeek LLM 67B better than o3?

o3 is the stronger model overall, scoring 47.5 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and o3 share?

15 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper