Model comparison

Devstral Small 2505 vs Gemini 2.5 Flash-Lite

Gemini 2.5 Flash-Lite is the stronger model overall, scoring 37.0 to 34.3 on the Noometry Index.

Last verified . 1 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Gemini 2.5 Flash-Lite Google

37.0

Rank #211 Confirmed

Summary

  • They share 1 benchmark with published results for both. Devstral Small 2505 scores higher in 1 category and Gemini 2.5 Flash-Lite in 1 category; one gap is clear of the uncertainty.
  • Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.10 / $0.40 for Gemini 2.5 Flash-Lite.
  • Gemini 2.5 Flash-Lite accepts more context: 1.05M tokens versus 128K.
  • Devstral Small 2505 has downloadable open weights; the other is API-only.

Side by side

Devstral Small 2505 and Gemini 2.5 Flash-Lite specifications
Devstral Small 2505Gemini 2.5 Flash-Lite
ProviderMistral AIGoogle
Noometry Index34.337.0
Released2025-05-072025-06-17
WeightsOpenProprietary
Context window128K1.05M
Max output128K66K
Input $ / M tokens$0.10$0.10
Output $ / M tokens$0.30$0.40
Results tracked433

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Devstral Small 2505: 38.9 (#166), Gemini 2.5 Flash-Lite: 38.5 (#173)

Coding benchmarks
BenchmarkDevstral Small 2505Gemini 2.5 Flash-Lite
SWE-bench Verified (bash only)56.4%—
SciCode28.8%—
WeirdML—35.2%
LMArena Coding—1373
ALE-Bench—325.9

Agentic & Tool Use Not comparable

Devstral Small 2505: —, Gemini 2.5 Flash-Lite: 28.0 (#96)

Agentic & Tool Use benchmarks
BenchmarkDevstral Small 2505Gemini 2.5 Flash-Lite
Berkeley Function Calling Leaderboard—36.9%

Reasoning Gemini 2.5 Flash-Lite leads

Devstral Small 2505: 19.7 (#252), Gemini 2.5 Flash-Lite: 22.2 (#205)

Reasoning benchmarks
BenchmarkDevstral Small 2505Gemini 2.5 Flash-Lite
Kagi LLM Benchmark37.7%40.5%
CritPt0%—
LMArena Hard Prompts—1377
DTBench—62.8%
LMCA—18.1%
Epoch Capabilities Index—133.94

Math Not comparable

Devstral Small 2505: —, Gemini 2.5 Flash-Lite: 38.0 (#144)

Math benchmarks
BenchmarkDevstral Small 2505Gemini 2.5 Flash-Lite
Omni-MATH—48%
LMArena Math—1373

Knowledge Not comparable

Devstral Small 2505: —, Gemini 2.5 Flash-Lite: 32.5 (#210)

Knowledge benchmarks
BenchmarkDevstral Small 2505Gemini 2.5 Flash-Lite
MMLU-Pro—53.7%
Vectara Hallucination Rate—3.3%
GPQA (HELM)—30.9%
LMArena Expert—1373

Multimodal Not comparable

Devstral Small 2505: —, Gemini 2.5 Flash-Lite: 29.1 (#114)

Multimodal benchmarks
BenchmarkDevstral Small 2505Gemini 2.5 Flash-Lite
LMArena Vision—1198
VPCT—30%

Multilingual Not comparable

Devstral Small 2505: —, Gemini 2.5 Flash-Lite: 49.3 (#134)

Multilingual benchmarks
BenchmarkDevstral Small 2505Gemini 2.5 Flash-Lite
LMArena Non-English—1369
LMArena Chinese—1404
LMArena French—1388
LMArena German—1389
LMArena Japanese—1359
LMArena Korean—1360
LMArena Russian—1373
LMArena Spanish—1396

Instruction Following Not comparable

Devstral Small 2505: —, Gemini 2.5 Flash-Lite: 70.0 (#168)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Gemini 2.5 Flash-Lite
IFEval—81%
LMArena Instruction Following—1367

Long Context Not comparable

Devstral Small 2505: —, Gemini 2.5 Flash-Lite: 33.3 (#262)

Long Context benchmarks
BenchmarkDevstral Small 2505Gemini 2.5 Flash-Lite
Fiction.LiveBench—47.2%
LMArena Longer Query—1373

Writing & Preference Not comparable

Devstral Small 2505: —, Gemini 2.5 Flash-Lite: 56.8 (#135)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Gemini 2.5 Flash-Lite
LMArena Text—1379
LMArena Creative Writing—1367
WildBench—81.8%
LMArena Multi-Turn—1366

Frequently asked questions

Is Devstral Small 2505 better than Gemini 2.5 Flash-Lite?

Gemini 2.5 Flash-Lite is the stronger model overall, scoring 37.0 to 34.3 on the Noometry Index.

Which is cheaper, Devstral Small 2505 or Gemini 2.5 Flash-Lite?

Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Gemini 2.5 Flash-Lite lists at $0.10 and $0.40.

Is Devstral Small 2505 or Gemini 2.5 Flash-Lite better for coding?

They score almost the same on coding (38.9 vs 38.5); test both on your own repository before choosing.

Which has the bigger context window?

Gemini 2.5 Flash-Lite does, with 1.05M tokens against 128K.

How many benchmarks do Devstral Small 2505 and Gemini 2.5 Flash-Lite share?

1 benchmark has published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Gemini 2.5 Flash-Lite has 33.

Related comparisons

Go deeper