Model comparison

Gemini 1.5 Flash (May 2024) vs Mistral Medium

Mistral Medium is the stronger model overall, scoring 36.3 to 33.2 on the Noometry Index.

Last verified . 24 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 3 categories and Mistral Medium in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 48.7.
  • The biggest single-benchmark swing is DTBench: 53.8% for Gemini 1.5 Flash (May 2024) and 75.5% for Mistral Medium.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash (May 2024) and Mistral Medium specifications
Gemini 1.5 Flash (May 2024)Mistral Medium
ProviderGoogleMistral AI
Noometry Index33.236.3
Released2024-05-142023-12-11
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked4236

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemini 1.5 Flash (May 2024): 34.4 (#236), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Medium
WeirdML24.9%43.7%
LMArena Coding12611434
FrontierCode—8%
SciCode—40.2%
BigCodeBench Instruct43.5%—
BigCodeBench Complete55.1%—
ALE-Bench—763.98
HumanEval+75.6%—
MBPP+67.5%—

Agentic & Tool Use Mistral Medium leads

Gemini 1.5 Flash (May 2024): 26.6 (#102), Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Medium
Berkeley Function Calling Leaderboard—37.7%
BALROG14.6%—

Reasoning Mistral Medium leads

Gemini 1.5 Flash (May 2024): 21.7 (#215), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Medium
LMArena Hard Prompts12571426
DTBench53.8%75.5%
Kagi LLM Benchmark—50%
CritPt—0%
LMCA—26.1%
Surface Evolver Bench—26.9%
Epoch Capabilities Index129.36—
ForecastBench53.9—
PIQA87.5%—

Math Mistral Medium leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Medium
OTIS Mock AIME 2024-202516.3%32.2%
LMArena Math12691408
MATH Level 561.9%81.6%
FrontierMath (Feb 2025 set)0%0.3%
ProofBench—9%
Omni-MATH30.4%—
GSM8K82.4%—

Knowledge Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 26.2 (#260), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Medium
GPQA Diamond47.3%59.5%
LMArena Expert12331408
Humanity's Last Exam—4.5%
MMLU-Pro67.8%—
Vectara Hallucination Rate—22.7%
GPQA (HELM)43.7%—
BoolQ85.8%—
MMLU77.9%—

Multimodal Too close to call

Gemini 1.5 Flash (May 2024): 36.0 (#81), Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Medium
LMArena Vision11411172
Video-MME70.3%—
GeoBench76%—

Multilingual Mistral Medium leads

Gemini 1.5 Flash (May 2024): 42.9 (#189), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Medium
LMArena Non-English12781408
LMArena Chinese12951447
LMArena French12581459
LMArena German12621432
LMArena Japanese12521378
LMArena Korean12211380
LMArena Russian12881411
LMArena Spanish12431433

Instruction Following Mistral Medium leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Medium
LMArena Instruction Following12581398
IFEval83.1%—

Long Context Mistral Medium leads

Gemini 1.5 Flash (May 2024): 39.0 (#187), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Medium
LMArena Longer Query12841406

Writing & Preference Mistral Medium leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Medium
LMArena Text12871424
LMArena Creative Writing12851391
LMArena Multi-Turn12531418
Short-Story Creative Writing—77.3%
WildBench79.2%—

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than Mistral Medium?

Mistral Medium is the stronger model overall, scoring 36.3 to 33.2 on the Noometry Index.

Is Gemini 1.5 Flash (May 2024) or Mistral Medium better for coding?

They score almost the same on coding (34.4 vs 34.2); test both on your own repository before choosing.

How many benchmarks do Gemini 1.5 Flash (May 2024) and Mistral Medium share?

24 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper