Model comparison

DeepSeek LLM 67B vs Gemini 1.5 Pro (May 2024)

Gemini 1.5 Pro (May 2024) is the stronger model overall, scoring 32.1 to 24.9 on the Noometry Index.

Last verified . 14 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Gemini 1.5 Pro (May 2024) Google

32.1

Rank #261 Confirmed

Summary

  • They share 14 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 1 category and Gemini 1.5 Pro (May 2024) in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Gemini 1.5 Pro (May 2024) leads 29.4 to 7.0.
  • The biggest single-benchmark swing is MATH Level 5: 6.4% for DeepSeek LLM 67B and 70.4% for Gemini 1.5 Pro (May 2024).
  • DeepSeek LLM 67B has downloadable open weights; the other is API-only.

Side by side

DeepSeek LLM 67B and Gemini 1.5 Pro (May 2024) specifications
DeepSeek LLM 67BGemini 1.5 Pro (May 2024)
ProviderDeepSeekGoogle
Noometry Index24.932.1
Released2023-11-292024-02-15
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1545

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Pro (May 2024) leads

DeepSeek LLM 67B: 31.9 (#278), Gemini 1.5 Pro (May 2024): 34.2 (#241)

Coding benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Pro (May 2024)
LMArena Coding10961294
WeirdML—22.2%
BigCodeBench Instruct—43.8%
BigCodeBench Complete—57.5%
CadEval—34%
HumanEval+—79.3%
MBPP+—74.6%

Agentic & Tool Use Not comparable

DeepSeek LLM 67B: —, Gemini 1.5 Pro (May 2024): 17.9 (#145)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Pro (May 2024)
TheAgentCompany—3.4%
Cybench—7.5%
BALROG—21%

Reasoning DeepSeek LLM 67B leads

DeepSeek LLM 67B: 16.5 (#304), Gemini 1.5 Pro (May 2024): 12.3 (#338)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Pro (May 2024)
LMArena Hard Prompts10701296
Epoch Capabilities Index110.5131.73
ARC-AGI-2—0.8%
SimpleBench—27.1%
Chess Puzzles0%—
DTBench—59%
BIG-Bench Hard—89.2%
ForecastBench—58.4

Math Gemini 1.5 Pro (May 2024) leads

DeepSeek LLM 67B: 8.7 (#324), Gemini 1.5 Pro (May 2024): 25.8 (#266)

Math benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Pro (May 2024)
OTIS Mock AIME 2024-20250.8%23.1%
LMArena Math11081315
MATH Level 56.4%70.4%
Omni-MATH—36.4%

Knowledge Gemini 1.5 Pro (May 2024) leads

DeepSeek LLM 67B: 7.0 (#313), Gemini 1.5 Pro (May 2024): 29.4 (#239)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Pro (May 2024)
GPQA Diamond24.6%57.2%
Humanity's Last Exam—4.6%
MMLU-Pro—73.7%
Confabulations—13.5%
GPQA (HELM)—53.4%
LMArena Expert—1279
MMLU—86.9%

Multimodal Not comparable

DeepSeek LLM 67B: —, Gemini 1.5 Pro (May 2024): 36.8 (#77)

Multimodal benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Pro (May 2024)
LMArena Vision—1161
Video-MME—75%

Multilingual Gemini 1.5 Pro (May 2024) leads

DeepSeek LLM 67B: 29.4 (#267), Gemini 1.5 Pro (May 2024): 45.3 (#174)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Pro (May 2024)
LMArena Non-English10731312
LMArena Chinese11321331
LMArena French—1302
LMArena German—1286
LMArena Japanese—1292
LMArena Korean—1298
LMArena Russian—1320
LMArena Spanish—1311

Instruction Following Gemini 1.5 Pro (May 2024) leads

DeepSeek LLM 67B: 55.4 (#277), Gemini 1.5 Pro (May 2024): 68.6 (#185)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Pro (May 2024)
LMArena Instruction Following10791297
IFEval—83.7%

Long Context Gemini 1.5 Pro (May 2024) leads

DeepSeek LLM 67B: 33.1 (#265), Gemini 1.5 Pro (May 2024): 39.8 (#169)

Long Context benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Pro (May 2024)
LMArena Longer Query10921308

Writing & Preference Gemini 1.5 Pro (May 2024) leads

DeepSeek LLM 67B: 31.6 (#282), Gemini 1.5 Pro (May 2024): 52.4 (#172)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Pro (May 2024)
LMArena Text11051319
LMArena Creative Writing10671333
LMArena Multi-Turn10821296
WildBench—81.3%

Frequently asked questions

Is DeepSeek LLM 67B better than Gemini 1.5 Pro (May 2024)?

Gemini 1.5 Pro (May 2024) is the stronger model overall, scoring 32.1 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or Gemini 1.5 Pro (May 2024) better for coding?

Gemini 1.5 Pro (May 2024) scores higher on coding benchmarks: 34.2 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and Gemini 1.5 Pro (May 2024) share?

14 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and Gemini 1.5 Pro (May 2024) has 45.

Related comparisons

Go deeper