Model comparison

Gemini 1.5 Flash (May 2024) vs Mistral Large 3

Mistral Large 3 is the stronger model overall, scoring 39.1 to 33.2 on the Noometry Index.

Last verified . 18 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 2 categories and Mistral Large 3 in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Large 3 leads 38.7 to 22.1.
  • Mistral Large 3 has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash (May 2024) and Mistral Large 3 specifications
Gemini 1.5 Flash (May 2024)Mistral Large 3
ProviderGoogleMistral AI
Noometry Index33.239.1
Released2024-05-142025-12-02
WeightsProprietaryOpen
Context window—262K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.75
Results tracked4224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemini 1.5 Flash (May 2024): 34.4 (#236), Mistral Large 3: 34.4 (#237)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large 3
LMArena Coding12611448
LMArena WebDev—1230
WeirdML24.9%—
BigCodeBench Instruct43.5%—
BigCodeBench Complete55.1%—
HumanEval+75.6%—
MBPP+67.5%—

Agentic & Tool Use Not comparable

Gemini 1.5 Flash (May 2024): 26.6 (#102), Mistral Large 3: —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large 3
BALROG14.6%—

Reasoning Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 21.7 (#215), Mistral Large 3: 15.2 (#319)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large 3
LMArena Hard Prompts12571429
Kagi LLM Benchmark—50.9%
NYT Connections (extended)—7.5%
Thematic Generalization—23%
DTBench53.8%—
Epoch Capabilities Index129.36—
ForecastBench53.9—
PIQA87.5%—

Math Mistral Large 3 leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), Mistral Large 3: 38.7 (#129)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large 3
LMArena Math12691414
OTIS Mock AIME 2024-202516.3%—
Omni-MATH30.4%—
MATH Level 561.9%—
FrontierMath (Feb 2025 set)0%—
GSM8K82.4%—

Knowledge Mistral Large 3 leads

Gemini 1.5 Flash (May 2024): 26.2 (#260), Mistral Large 3: 36.0 (#177)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large 3
LMArena Expert12331421
GPQA Diamond47.3%—
MMLU-Pro67.8%—
Vectara Hallucination Rate—14.5%
GPQA (HELM)43.7%—
BoolQ85.8%—
MMLU77.9%—

Multimodal Mistral Large 3 leads

Gemini 1.5 Flash (May 2024): 36.0 (#81), Mistral Large 3: 38.2 (#66)

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large 3
LMArena Vision11411221
Video-MME70.3%—
GeoBench76%—

Multilingual Mistral Large 3 leads

Gemini 1.5 Flash (May 2024): 42.9 (#189), Mistral Large 3: 52.5 (#84)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large 3
LMArena Non-English12781413
LMArena Chinese12951447
LMArena French12581455
LMArena German12621437
LMArena Japanese12521394
LMArena Korean12211384
LMArena Russian12881411
LMArena Spanish12431440

Instruction Following Mistral Large 3 leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), Mistral Large 3: 74.0 (#108)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large 3
LMArena Instruction Following12581403
IFEval83.1%—

Long Context Mistral Large 3 leads

Gemini 1.5 Flash (May 2024): 39.0 (#187), Mistral Large 3: 43.1 (#105)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large 3
LMArena Longer Query12841413

Writing & Preference Mistral Large 3 leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), Mistral Large 3: 60.0 (#101)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Mistral Large 3
LMArena Text12871428
LMArena Creative Writing12851386
LMArena Multi-Turn12531429
EQ-Bench Creative Writing—1412
WildBench79.2%—

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than Mistral Large 3?

Mistral Large 3 is the stronger model overall, scoring 39.1 to 33.2 on the Noometry Index.

Is Gemini 1.5 Flash (May 2024) or Mistral Large 3 better for coding?

They score almost the same on coding (34.4 vs 34.4); test both on your own repository before choosing.

How many benchmarks do Gemini 1.5 Flash (May 2024) and Mistral Large 3 share?

18 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and Mistral Large 3 has 24.

Related comparisons

Go deeper