Model comparison

Gemini 1.5 Flash (May 2024) vs Mistral Large

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 31.9 on the Noometry Index.

Last verified . 34 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 34 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 6 categories and Mistral Large in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemini 1.5 Flash (May 2024) leads 48.7 to 40.7.
  • The biggest single-benchmark swing is BigCodeBench Complete: 55.1% for Gemini 1.5 Flash (May 2024) and 38.3% for Mistral Large.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash (May 2024) and Mistral Large specifications
Gemini 1.5 Flash (May 2024)Mistral Large
ProviderGoogleMistral AI
Noometry Index33.231.9
Released2024-05-142024-02-26
WeightsProprietaryOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked4251

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemini 1.5 Flash (May 2024): 34.4 (#236), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large
BigCodeBench Instruct43.5%30%
LMArena Coding12611277
BigCodeBench Complete55.1%38.3%
HumanEval+75.6%62.2%
MBPP+67.5%59.5%
SciCode—36.2%
WeirdML24.9%—
LiveBench Coding—47.1%
ALE-Bench—264.7

Agentic & Tool Use Mistral Large leads

Gemini 1.5 Flash (May 2024): 26.6 (#102), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large
Berkeley Function Calling Leaderboard—38.4%
BALROG14.6%—

Reasoning Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 21.7 (#215), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large
LMArena Hard Prompts12571257
DTBench53.8%65.1%
Epoch Capabilities Index129.36128.52
ForecastBench53.957.1
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
LiveBench Data Analysis—50.1%
LMCA—16.7%
LiveBench—48.4%
PIQA87.5%—

Math Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large
OTIS Mock AIME 2024-202516.3%8.5%
Omni-MATH30.4%28.1%
LMArena Math12691262
MATH Level 561.9%50.3%
FrontierMath (Feb 2025 set)0%0.3%
LiveBench Math—42.5%
GSM8K82.4%—

Knowledge Mistral Large leads

Gemini 1.5 Flash (May 2024): 26.2 (#260), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large
GPQA Diamond47.3%51.3%
MMLU-Pro67.8%59.9%
GPQA (HELM)43.7%43.5%
LMArena Expert12331232
MMLU77.9%80%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
BoolQ85.8%—

Multimodal Not comparable

Gemini 1.5 Flash (May 2024): 36.0 (#81), Mistral Large: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large
LMArena Vision1141—
Video-MME70.3%—
GeoBench76%—

Multilingual Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 42.9 (#189), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large
LMArena Non-English12781237
LMArena Chinese12951240
LMArena French12581325
LMArena German12621254
LMArena Japanese12521188
LMArena Korean12211202
LMArena Russian12881257
LMArena Spanish12431268

Instruction Following Mistral Large leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large
IFEval83.1%87.7%
LMArena Instruction Following12581249
LiveBench Instruction Following—67.9%

Long Context Too close to call

Gemini 1.5 Flash (May 2024): 39.0 (#187), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large
LMArena Longer Query12841261

Writing & Preference Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large
LMArena Text12871266
LMArena Creative Writing12851243
WildBench79.2%80.1%
LMArena Multi-Turn12531260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
LiveBench Language—39.4%

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than Mistral Large?

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 31.9 on the Noometry Index.

Is Gemini 1.5 Flash (May 2024) or Mistral Large better for coding?

They score almost the same on coding (34.4 vs 34.3); test both on your own repository before choosing.

How many benchmarks do Gemini 1.5 Flash (May 2024) and Mistral Large share?

34 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper