Model comparison

DeepSeek LLM 67B vs o1

o1 is the stronger model overall, scoring 40.9 to 24.9 on the Noometry Index.

Last verified . 15 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 15 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 0 categories and o1 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o1 leads 41.5 to 7.0.
  • The biggest single-benchmark swing is MATH Level 5: 6.4% for DeepSeek LLM 67B and 94.7% for o1.
  • DeepSeek LLM 67B has downloadable open weights; the other is API-only.

Side by side

DeepSeek LLM 67B and o1 specifications
DeepSeek LLM 67Bo1
ProviderDeepSeekOpenAI
Noometry Index24.940.9
Released2023-11-292024-09-12
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked1552

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

DeepSeek LLM 67B: 31.9 (#278), o1: 46.1 (#70)

Coding benchmarks
BenchmarkDeepSeek LLM 67Bo1
LMArena Coding10961367
Aider Polyglot—61.7%
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

DeepSeek LLM 67B: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek LLM 67Bo1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

DeepSeek LLM 67B: 16.5 (#304), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67Bo1
Chess Puzzles0%15%
LMArena Hard Prompts10701371
Epoch Capabilities Index110.5141.91
SimpleBench—41.7%
ARC-AGI-1—30.7%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
LiveBench—75.7%

Math o1 leads

DeepSeek LLM 67B: 8.7 (#324), o1: 36.1 (#175)

Math benchmarks
BenchmarkDeepSeek LLM 67Bo1
OTIS Mock AIME 2024-20250.8%73.3%
LMArena Math11081388
MATH Level 56.4%94.7%
FrontierMath (Tiers 1-3)—14.7%
LiveBench Math—80.3%
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

DeepSeek LLM 67B: 7.0 (#313), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67Bo1
GPQA Diamond24.6%76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%
LMArena Expert—1361

Multimodal Not comparable

DeepSeek LLM 67B: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkDeepSeek LLM 67Bo1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual o1 leads

DeepSeek LLM 67B: 29.4 (#267), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67Bo1
LMArena Non-English10731358
LMArena Chinese11321394
LMArena French—1344
LMArena German—1337
LMArena Japanese—1346
LMArena Korean—1396
LMArena Russian—1356
LMArena Spanish—1345

Instruction Following o1 leads

DeepSeek LLM 67B: 55.4 (#277), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67Bo1
LMArena Instruction Following10791367
LiveBench Instruction Following—81.5%

Long Context o1 leads

DeepSeek LLM 67B: 33.1 (#265), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkDeepSeek LLM 67Bo1
LMArena Longer Query10921378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

DeepSeek LLM 67B: 31.6 (#282), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67Bo1
LMArena Text11051366
LMArena Creative Writing10671348
LMArena Multi-Turn10821369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is DeepSeek LLM 67B better than o1?

o1 is the stronger model overall, scoring 40.9 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and o1 share?

15 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper