Model comparison

Gemini 2.5 Flash vs Mistral Medium 3.1

Gemini 2.5 Flash is the stronger model overall, scoring 39.3 to 31.9 on the Noometry Index.

Last verified . 1 shared benchmarks.

Gemini 2.5 Flash Google

39.3

Rank #170 Confirmed

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Summary

  • They share 1 benchmark with published results for both. Gemini 2.5 Flash scores higher in 1 category and Mistral Medium 3.1 in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 2.5 Flash leads 18.1 to 10.6.
  • Mistral Medium 3.1 is cheaper at $0.40 / $2 per million input/output tokens, against $0.30 / $2.50 for Gemini 2.5 Flash.
  • Gemini 2.5 Flash accepts more context: 1.05M tokens versus 131K.

Side by side

Gemini 2.5 Flash and Mistral Medium 3.1 specifications
Gemini 2.5 FlashMistral Medium 3.1
ProviderGoogleMistral AI
Noometry Index39.331.9
Released2025-04-17—
WeightsProprietaryProprietary
Context window1.05M131K
Max output66K105K
Input $ / M tokens$0.30$0.40
Output $ / M tokens$2.50$2
Results tracked543

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Gemini 2.5 Flash: 35.8 (#220), Mistral Medium 3.1: —

Coding benchmarks
BenchmarkGemini 2.5 FlashMistral Medium 3.1
SWE-bench Verified (bash only)28.7%—
Aider Polyglot55.1%—
WeirdML41.9%—
LMArena Coding1424—
ALE-Bench661.88—

Agentic & Tool Use Not comparable

Gemini 2.5 Flash: 30.8 (#74), Mistral Medium 3.1: —

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 FlashMistral Medium 3.1
Terminal-Bench17.1%—
Berkeley Function Calling Leaderboard56.2%—
TheAgentCompany41.1%—
BALROG33.5%—
Vending-Bench 2548.84—

Reasoning Gemini 2.5 Flash leads

Gemini 2.5 Flash: 18.1 (#286), Mistral Medium 3.1: 10.6 (#341)

Reasoning benchmarks
BenchmarkGemini 2.5 FlashMistral Medium 3.1
ARC-AGI-22.5%—
SimpleBench41.2%—
Kagi LLM Benchmark56.8%—
NYT Connections (extended)—6.5%
ARC-AGI-133.3%—
CritPt1.1%—
EnigmaEval2.7%—
Thematic Generalization—20.3%
LMArena Hard Prompts1422—
DTBench76.5%—
LMCA27.5%—
Epoch Capabilities Index143.03—
ForecastBench60.6—

Math Not comparable

Gemini 2.5 Flash: 39.9 (#98), Mistral Medium 3.1: —

Math benchmarks
BenchmarkGemini 2.5 FlashMistral Medium 3.1
OTIS Mock AIME 2024-202573.1%—
Omni-MATH38.5%—
LMArena Math1415—
FrontierMath (Feb 2025 set)4.8%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Not comparable

Gemini 2.5 Flash: 36.4 (#168), Mistral Medium 3.1: —

Knowledge benchmarks
BenchmarkGemini 2.5 FlashMistral Medium 3.1
Humanity's Last Exam12.1%—
MMLU-Pro63.9%—
Confabulations16.8%—
Vectara Hallucination Rate7.8%—
GPQA (HELM)39%—
LMArena Expert1426—

Multimodal Not comparable

Gemini 2.5 Flash: 41.8 (#32), Mistral Medium 3.1: —

Multimodal benchmarks
BenchmarkGemini 2.5 FlashMistral Medium 3.1
LMArena Vision1253—
GeoBench76%—
VPCT46.2%—
SpatialViz-Bench36.9%—

Multilingual Not comparable

Gemini 2.5 Flash: 52.3 (#88), Mistral Medium 3.1: —

Multilingual benchmarks
BenchmarkGemini 2.5 FlashMistral Medium 3.1
LMArena Non-English1409—
LMArena Chinese1450—
LMArena French1433—
LMArena German1418—
LMArena Japanese1405—
LMArena Korean1385—
LMArena Russian1415—
LMArena Spanish1421—

Instruction Following Not comparable

Gemini 2.5 Flash: 75.7 (#54), Mistral Medium 3.1: —

Instruction Following benchmarks
BenchmarkGemini 2.5 FlashMistral Medium 3.1
IFEval89.8%—
LMArena Instruction Following1405—

Long Context Not comparable

Gemini 2.5 Flash: 47.5 (#17), Mistral Medium 3.1: —

Long Context benchmarks
BenchmarkGemini 2.5 FlashMistral Medium 3.1
Fiction.LiveBench77.8%—
LMArena Longer Query1419—

Writing & Preference Mistral Medium 3.1 leads

Gemini 2.5 Flash: 53.8 (#157), Mistral Medium 3.1: 55.5 (#145)

Writing & Preference benchmarks
BenchmarkGemini 2.5 FlashMistral Medium 3.1
EQ-Bench Creative Writing11371476
LMArena Text1417—
LMArena Creative Writing1400—
Short-Story Creative Writing76.5%—
WildBench81.7%—
LMArena Multi-Turn1408—

Frequently asked questions

Is Gemini 2.5 Flash better than Mistral Medium 3.1?

Gemini 2.5 Flash is the stronger model overall, scoring 39.3 to 31.9 on the Noometry Index.

Which is cheaper, Gemini 2.5 Flash or Mistral Medium 3.1?

Mistral Medium 3.1 is cheaper. It lists at $0.40 per million input tokens and $2 per million output tokens; Gemini 2.5 Flash lists at $0.30 and $2.50.

Which has the bigger context window?

Gemini 2.5 Flash does, with 1.05M tokens against 131K.

How many benchmarks do Gemini 2.5 Flash and Mistral Medium 3.1 share?

1 benchmark has published results for both models. Gemini 2.5 Flash has 54 scored results on Noometry and Mistral Medium 3.1 has 3.

Related comparisons

Go deeper