Model comparison

Gemini 1.5 Pro (May 2024) vs Mistral Large 3

Mistral Large 3 is the stronger model overall, scoring 39.1 to 32.1 on the Noometry Index.

Last verified . 18 shared benchmarks.

Gemini 1.5 Pro (May 2024) Google

32.1

Rank #261 Confirmed

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Gemini 1.5 Pro (May 2024) scores higher in 0 categories and Mistral Large 3 in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Large 3 leads 38.7 to 25.8.
  • Mistral Large 3 has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Pro (May 2024) and Mistral Large 3 specifications
Gemini 1.5 Pro (May 2024)Mistral Large 3
ProviderGoogleMistral AI
Noometry Index32.139.1
Released2024-02-152025-12-02
WeightsProprietaryOpen
Context window—262K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.75
Results tracked4524

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemini 1.5 Pro (May 2024): 34.2 (#241), Mistral Large 3: 34.4 (#237)

Coding benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Mistral Large 3
LMArena Coding12941448
LMArena WebDev—1230
WeirdML22.2%—
BigCodeBench Instruct43.8%—
BigCodeBench Complete57.5%—
CadEval34%—
HumanEval+79.3%—
MBPP+74.6%—

Agentic & Tool Use Not comparable

Gemini 1.5 Pro (May 2024): 17.9 (#145), Mistral Large 3: —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Mistral Large 3
TheAgentCompany3.4%—
Cybench7.5%—
BALROG21%—

Reasoning Mistral Large 3 leads

Gemini 1.5 Pro (May 2024): 12.3 (#338), Mistral Large 3: 15.2 (#319)

Reasoning benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Mistral Large 3
LMArena Hard Prompts12961429
ARC-AGI-20.8%—
SimpleBench27.1%—
Kagi LLM Benchmark—50.9%
NYT Connections (extended)—7.5%
Thematic Generalization—23%
DTBench59%—
BIG-Bench Hard89.2%—
Epoch Capabilities Index131.73—
ForecastBench58.4—

Math Mistral Large 3 leads

Gemini 1.5 Pro (May 2024): 25.8 (#266), Mistral Large 3: 38.7 (#129)

Math benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Mistral Large 3
LMArena Math13151414
OTIS Mock AIME 2024-202523.1%—
Omni-MATH36.4%—
MATH Level 570.4%—

Knowledge Mistral Large 3 leads

Gemini 1.5 Pro (May 2024): 29.4 (#239), Mistral Large 3: 36.0 (#177)

Knowledge benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Mistral Large 3
LMArena Expert12791421
GPQA Diamond57.2%—
Humanity's Last Exam4.6%—
MMLU-Pro73.7%—
Confabulations13.5%—
Vectara Hallucination Rate—14.5%
GPQA (HELM)53.4%—
MMLU86.9%—

Multimodal Mistral Large 3 leads

Gemini 1.5 Pro (May 2024): 36.8 (#77), Mistral Large 3: 38.2 (#66)

Multimodal benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Mistral Large 3
LMArena Vision11611221
Video-MME75%—

Multilingual Mistral Large 3 leads

Gemini 1.5 Pro (May 2024): 45.3 (#174), Mistral Large 3: 52.5 (#84)

Multilingual benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Mistral Large 3
LMArena Non-English13121413
LMArena Chinese13311447
LMArena French13021455
LMArena German12861437
LMArena Japanese12921394
LMArena Korean12981384
LMArena Russian13201411
LMArena Spanish13111440

Instruction Following Mistral Large 3 leads

Gemini 1.5 Pro (May 2024): 68.6 (#185), Mistral Large 3: 74.0 (#108)

Instruction Following benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Mistral Large 3
LMArena Instruction Following12971403
IFEval83.7%—

Long Context Mistral Large 3 leads

Gemini 1.5 Pro (May 2024): 39.8 (#169), Mistral Large 3: 43.1 (#105)

Long Context benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Mistral Large 3
LMArena Longer Query13081413

Writing & Preference Mistral Large 3 leads

Gemini 1.5 Pro (May 2024): 52.4 (#172), Mistral Large 3: 60.0 (#101)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Mistral Large 3
LMArena Text13191428
LMArena Creative Writing13331386
LMArena Multi-Turn12961429
EQ-Bench Creative Writing—1412
WildBench81.3%—

Frequently asked questions

Is Gemini 1.5 Pro (May 2024) better than Mistral Large 3?

Mistral Large 3 is the stronger model overall, scoring 39.1 to 32.1 on the Noometry Index.

Is Gemini 1.5 Pro (May 2024) or Mistral Large 3 better for coding?

They score almost the same on coding (34.2 vs 34.4); test both on your own repository before choosing.

How many benchmarks do Gemini 1.5 Pro (May 2024) and Mistral Large 3 share?

18 benchmarks have published results for both models. Gemini 1.5 Pro (May 2024) has 45 scored results on Noometry and Mistral Large 3 has 24.

Related comparisons

Go deeper