Model comparison

DeepSeek V4.1 Flash vs Gemini 2.5 Flash-Lite

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 37.0 on the Noometry Index.

Last verified . 22 shared benchmarks.

DeepSeek V4.1 Flash DeepSeek

52.8

Rank #38 Confirmed

Gemini 2.5 Flash-Lite Google

37.0

Rank #211 Confirmed

Summary

  • They share 22 benchmarks with published results for both. DeepSeek V4.1 Flash scores higher in 10 categories and Gemini 2.5 Flash-Lite in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek V4.1 Flash leads 66.7 to 38.0.
  • The biggest single-benchmark swing is LMCA: 47% for DeepSeek V4.1 Flash and 18.1% for Gemini 2.5 Flash-Lite.
  • Gemini 2.5 Flash-Lite is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.15 / $0.60 for DeepSeek V4.1 Flash.
  • Gemini 2.5 Flash-Lite accepts more context: 1.05M tokens versus 1M.
  • DeepSeek V4.1 Flash has downloadable open weights; the other is API-only.

Side by side

DeepSeek V4.1 Flash and Gemini 2.5 Flash-Lite specifications
DeepSeek V4.1 FlashGemini 2.5 Flash-Lite
ProviderDeepSeekGoogle
Noometry Index52.837.0
Released2026-09-092025-06-17
WeightsOpenProprietary
Context window1M1.05M
Max output393K66K
Input $ / M tokens$0.15$0.10
Output $ / M tokens$0.60$0.40
Results tracked3733

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 52.9 (#32), Gemini 2.5 Flash-Lite: 38.5 (#173)

Coding benchmarks
BenchmarkDeepSeek V4.1 FlashGemini 2.5 Flash-Lite
LMArena Coding15061373
ALE-Bench1,092325.9
LMArena WebDev1619—
SciCode51.9%—
WeirdML—35.2%

Agentic & Tool Use DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 31.2 (#69), Gemini 2.5 Flash-Lite: 28.0 (#96)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek V4.1 FlashGemini 2.5 Flash-Lite
APEX-Agents39.5%—
Berkeley Function Calling Leaderboard—36.9%
GDP.pdf19.8%—

Reasoning DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 50.2 (#36), Gemini 2.5 Flash-Lite: 22.2 (#205)

Reasoning benchmarks
BenchmarkDeepSeek V4.1 FlashGemini 2.5 Flash-Lite
LMArena Hard Prompts14831377
DTBench89.9%62.8%
LMCA47%18.1%
Epoch Capabilities Index154.9133.94
Kagi LLM Benchmark—40.5%
NYT Connections (extended)89.6%—
CritPt14.3%—
Mystery Game Puzzles43%—
Surface Evolver Bench46.3%—

Math DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 66.7 (#25), Gemini 2.5 Flash-Lite: 38.0 (#144)

Math benchmarks
BenchmarkDeepSeek V4.1 FlashGemini 2.5 Flash-Lite
LMArena Math14771373
FrontierMath (Tiers 1-3)67.4%—
FrontierMath Tier 426.8%—
OTIS Mock AIME 2024-202598.3%—
ProofBench54%—
Omni-MATH—48%

Knowledge DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 57.9 (#38), Gemini 2.5 Flash-Lite: 32.5 (#210)

Knowledge benchmarks
BenchmarkDeepSeek V4.1 FlashGemini 2.5 Flash-Lite
LMArena Expert15061373
GPQA Diamond89.8%—
MMLU-Pro—53.7%
Vectara Hallucination Rate—3.3%
GPQA (HELM)—30.9%

Multimodal DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 39.1 (#61), Gemini 2.5 Flash-Lite: 29.1 (#114)

Multimodal benchmarks
BenchmarkDeepSeek V4.1 FlashGemini 2.5 Flash-Lite
LMArena Vision12771198
VPCT—30%
Furniture Assembly34.2%—

Multilingual DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 55.0 (#35), Gemini 2.5 Flash-Lite: 49.3 (#134)

Multilingual benchmarks
BenchmarkDeepSeek V4.1 FlashGemini 2.5 Flash-Lite
LMArena Non-English14481369
LMArena Chinese14971404
LMArena French14521388
LMArena German14841389
LMArena Japanese14121359
LMArena Korean14521360
LMArena Russian14711373
LMArena Spanish14591396

Instruction Following DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 77.3 (#26), Gemini 2.5 Flash-Lite: 70.0 (#168)

Instruction Following benchmarks
BenchmarkDeepSeek V4.1 FlashGemini 2.5 Flash-Lite
LMArena Instruction Following14741367
IFEval—81%

Long Context DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 45.2 (#47), Gemini 2.5 Flash-Lite: 33.3 (#262)

Long Context benchmarks
BenchmarkDeepSeek V4.1 FlashGemini 2.5 Flash-Lite
LMArena Longer Query14751373
Fiction.LiveBench—47.2%

Writing & Preference DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 65.4 (#48), Gemini 2.5 Flash-Lite: 56.8 (#135)

Writing & Preference benchmarks
BenchmarkDeepSeek V4.1 FlashGemini 2.5 Flash-Lite
LMArena Text14621379
LMArena Creative Writing14351367
LMArena Multi-Turn14571366
EQ-Bench Creative Writing1540—
WildBench—81.8%

Frequently asked questions

Is DeepSeek V4.1 Flash better than Gemini 2.5 Flash-Lite?

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 37.0 on the Noometry Index.

Which is cheaper, DeepSeek V4.1 Flash or Gemini 2.5 Flash-Lite?

Gemini 2.5 Flash-Lite is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; DeepSeek V4.1 Flash lists at $0.15 and $0.60.

Is DeepSeek V4.1 Flash or Gemini 2.5 Flash-Lite better for coding?

DeepSeek V4.1 Flash scores higher on coding benchmarks: 52.9 versus 38.5 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash-Lite does, with 1.05M tokens against 1M.

How many benchmarks do DeepSeek V4.1 Flash and Gemini 2.5 Flash-Lite share?

22 benchmarks have published results for both models. DeepSeek V4.1 Flash has 37 scored results on Noometry and Gemini 2.5 Flash-Lite has 33.

Related comparisons

Go deeper