Model comparison

Gemini 2.0 Flash (Feb 2025) vs Mistral Large

Gemini 2.0 Flash (Feb 2025) is the stronger model overall, scoring 35.1 to 31.9 on the Noometry Index.

Last verified . 42 shared benchmarks.

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 42 benchmarks with published results for both. Gemini 2.0 Flash (Feb 2025) scores higher in 5 categories and Mistral Large in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemini 2.0 Flash (Feb 2025) leads 37.9 to 18.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 57.8% for Gemini 2.0 Flash (Feb 2025) and 8.5% for Mistral Large.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Gemini 2.0 Flash (Feb 2025) and Mistral Large specifications
Gemini 2.0 Flash (Feb 2025)Mistral Large
ProviderGoogleMistral AI
Noometry Index35.131.9
Released2024-12-062024-02-26
WeightsProprietaryOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked5451

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Gemini 2.0 Flash (Feb 2025): 28.4 (#315), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large
BigCodeBench Instruct45.9%30%
LiveBench Coding63.4%47.1%
LMArena Coding13501277
BigCodeBench Complete59.9%38.3%
SWE-bench Verified (bash only)13.5%—
Aider Polyglot38.2%—
SciCode—36.2%
WeirdML25.8%—
CadEval30%—
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Too close to call

Gemini 2.0 Flash (Feb 2025): 28.1 (#92), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large
Berkeley Function Calling Leaderboard—38.4%
TheAgentCompany11.4%—

Reasoning Too close to call

Gemini 2.0 Flash (Feb 2025): 15.2 (#318), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large
SimpleBench31.1%22.5%
LiveBench Reasoning78.2%43.5%
LMArena Hard Prompts13461257
DTBench63.2%65.1%
LiveBench Data Analysis69.4%50.1%
Epoch Capabilities Index135.36128.52
LiveBench66.9%48.4%
ARC-AGI-21.3%—
Kagi LLM Benchmark37.8%—
CritPt—0%
EnigmaEval1.1%—
LMCA—16.7%
ForecastBench—57.1

Math Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 37.9 (#146), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large
OTIS Mock AIME 2024-202557.8%8.5%
Omni-MATH45.9%28.1%
LiveBench Math75.8%42.5%
LMArena Math13521262
MATH Level 582.2%50.3%
FrontierMath (Feb 2025 set)1.7%0.3%

Knowledge Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 32.0 (#213), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large
GPQA Diamond64.1%51.3%
MMLU-Pro73.7%59.9%
Confabulations12.4%21.4%
GPQA (HELM)55.6%43.5%
LMArena Expert13391232
MMLU79.7%80%
Humanity's Last Exam6.6%—
Vectara Hallucination Rate—4.5%

Multimodal Not comparable

Gemini 2.0 Flash (Feb 2025): 36.5 (#79), Mistral Large: —

Multimodal benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large
LMArena Vision1158—
GeoBench77%—

Multilingual Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 47.4 (#149), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large
LMArena Non-English13421237
LMArena Chinese13731240
LMArena French13911325
LMArena German13531254
LMArena Japanese12941188
LMArena Korean13131202
LMArena Russian13511257
LMArena Spanish13631268

Instruction Following Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 74.4 (#97), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large
LiveBench Instruction Following85.8%67.9%
IFEval84.1%87.7%
LMArena Instruction Following13361249

Long Context Too close to call

Gemini 2.0 Flash (Feb 2025): 38.1 (#203), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large
LMArena Longer Query13441261
Fiction.LiveBench61.1%—

Writing & Preference Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 49.5 (#190), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large
LMArena Text13541266
LMArena Creative Writing13401243
Short-Story Creative Writing73.8%69%
EQ-Bench Creative Writing1128985
WildBench80%80.1%
LMArena Multi-Turn13501260
LiveBench Language51.3%39.4%

Frequently asked questions

Is Gemini 2.0 Flash (Feb 2025) better than Mistral Large?

Gemini 2.0 Flash (Feb 2025) is the stronger model overall, scoring 35.1 to 31.9 on the Noometry Index.

Is Gemini 2.0 Flash (Feb 2025) or Mistral Large better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 28.4 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Flash (Feb 2025) and Mistral Large share?

42 benchmarks have published results for both models. Gemini 2.0 Flash (Feb 2025) has 54 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper