Model comparison

DeepSeek LLM 67B vs Gemini 1.5 Flash (May 2024)

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 24.9 on the Noometry Index.

Last verified . 14 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Summary

  • They share 14 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 0 categories and Gemini 1.5 Flash (May 2024) in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Gemini 1.5 Flash (May 2024) leads 26.2 to 7.0.
  • The biggest single-benchmark swing is MATH Level 5: 6.4% for DeepSeek LLM 67B and 61.9% for Gemini 1.5 Flash (May 2024).
  • DeepSeek LLM 67B has downloadable open weights; the other is API-only.

Side by side

DeepSeek LLM 67B and Gemini 1.5 Flash (May 2024) specifications
DeepSeek LLM 67BGemini 1.5 Flash (May 2024)
ProviderDeepSeekGoogle
Noometry Index24.933.2
Released2023-11-292024-05-14
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1542

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash (May 2024) leads

DeepSeek LLM 67B: 31.9 (#278), Gemini 1.5 Flash (May 2024): 34.4 (#236)

Coding benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Flash (May 2024)
LMArena Coding10961261
WeirdML—24.9%
BigCodeBench Instruct—43.5%
BigCodeBench Complete—55.1%
HumanEval+—75.6%
MBPP+—67.5%

Agentic & Tool Use Not comparable

DeepSeek LLM 67B: —, Gemini 1.5 Flash (May 2024): 26.6 (#102)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Flash (May 2024)
BALROG—14.6%

Reasoning Gemini 1.5 Flash (May 2024) leads

DeepSeek LLM 67B: 16.5 (#304), Gemini 1.5 Flash (May 2024): 21.7 (#215)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Flash (May 2024)
LMArena Hard Prompts10701257
Epoch Capabilities Index110.5129.36
Chess Puzzles0%—
DTBench—53.8%
ForecastBench—53.9
PIQA—87.5%

Math Gemini 1.5 Flash (May 2024) leads

DeepSeek LLM 67B: 8.7 (#324), Gemini 1.5 Flash (May 2024): 22.1 (#281)

Math benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Flash (May 2024)
OTIS Mock AIME 2024-20250.8%16.3%
LMArena Math11081269
MATH Level 56.4%61.9%
Omni-MATH—30.4%
FrontierMath (Feb 2025 set)—0%
GSM8K—82.4%

Knowledge Gemini 1.5 Flash (May 2024) leads

DeepSeek LLM 67B: 7.0 (#313), Gemini 1.5 Flash (May 2024): 26.2 (#260)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Flash (May 2024)
GPQA Diamond24.6%47.3%
MMLU-Pro—67.8%
GPQA (HELM)—43.7%
LMArena Expert—1233
BoolQ—85.8%
MMLU—77.9%

Multimodal Not comparable

DeepSeek LLM 67B: —, Gemini 1.5 Flash (May 2024): 36.0 (#81)

Multimodal benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Flash (May 2024)
LMArena Vision—1141
Video-MME—70.3%
GeoBench—76%

Multilingual Gemini 1.5 Flash (May 2024) leads

DeepSeek LLM 67B: 29.4 (#267), Gemini 1.5 Flash (May 2024): 42.9 (#189)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Flash (May 2024)
LMArena Non-English10731278
LMArena Chinese11321295
LMArena French—1258
LMArena German—1262
LMArena Japanese—1252
LMArena Korean—1221
LMArena Russian—1288
LMArena Spanish—1243

Instruction Following Gemini 1.5 Flash (May 2024) leads

DeepSeek LLM 67B: 55.4 (#277), Gemini 1.5 Flash (May 2024): 66.8 (#205)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Flash (May 2024)
LMArena Instruction Following10791258
IFEval—83.1%

Long Context Gemini 1.5 Flash (May 2024) leads

DeepSeek LLM 67B: 33.1 (#265), Gemini 1.5 Flash (May 2024): 39.0 (#187)

Long Context benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Flash (May 2024)
LMArena Longer Query10921284

Writing & Preference Gemini 1.5 Flash (May 2024) leads

DeepSeek LLM 67B: 31.6 (#282), Gemini 1.5 Flash (May 2024): 48.7 (#196)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BGemini 1.5 Flash (May 2024)
LMArena Text11051287
LMArena Creative Writing10671285
LMArena Multi-Turn10821253
WildBench—79.2%

Frequently asked questions

Is DeepSeek LLM 67B better than Gemini 1.5 Flash (May 2024)?

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or Gemini 1.5 Flash (May 2024) better for coding?

Gemini 1.5 Flash (May 2024) scores higher on coding benchmarks: 34.4 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and Gemini 1.5 Flash (May 2024) share?

14 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and Gemini 1.5 Flash (May 2024) has 42.

Related comparisons

Go deeper