Model comparison

Gemini 2.0 Flash (Feb 2025) vs Mistral Large 4

Mistral Large 4 is the stronger model overall, scoring 43.1 to 35.1 on the Noometry Index.

Last verified . 12 shared benchmarks.

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Mistral Large 4 Mistral AI

43.1

Rank #99 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Gemini 2.0 Flash (Feb 2025) scores higher in 0 categories and Mistral Large 4 in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Mistral Large 4 leads 48.6 to 28.4.

Side by side

Gemini 2.0 Flash (Feb 2025) and Mistral Large 4 specifications
Gemini 2.0 Flash (Feb 2025)Mistral Large 4
ProviderGoogleMistral AI
Noometry Index35.143.1
Released2024-12-062026-10-06
WeightsProprietaryProprietary
Context window—1.05M
Max output—262K
Input $ / M tokens—$0.68
Output $ / M tokens—$2.09
Results tracked5415

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large 4 leads

Gemini 2.0 Flash (Feb 2025): 28.4 (#315), Mistral Large 4: 48.6 (#57)

Coding benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large 4
LMArena Coding13501475
SWE-bench Verified (bash only)13.5%—
Aider Polyglot38.2%—
LMArena WebDev—1541
WeirdML25.8%—
BigCodeBench Instruct45.9%—
LiveBench Coding63.4%—
BigCodeBench Complete59.9%—
CadEval30%—

Agentic & Tool Use Not comparable

Gemini 2.0 Flash (Feb 2025): 28.1 (#92), Mistral Large 4: —

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large 4
TheAgentCompany11.4%—

Reasoning Mistral Large 4 leads

Gemini 2.0 Flash (Feb 2025): 15.2 (#318), Mistral Large 4: 22.5 (#192)

Reasoning benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large 4
LMArena Hard Prompts13461444
ARC-AGI-21.3%—
SimpleBench31.1%—
Kagi LLM Benchmark37.8%—
NYT Connections (extended)—27.4%
EnigmaEval1.1%—
LiveBench Reasoning78.2%—
DTBench63.2%—
LiveBench Data Analysis69.4%—
Epoch Capabilities Index135.36—
LiveBench66.9%—

Math Mistral Large 4 leads

Gemini 2.0 Flash (Feb 2025): 37.9 (#146), Mistral Large 4: 40.4 (#91)

Math benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large 4
LMArena Math13521488
OTIS Mock AIME 2024-202557.8%—
Omni-MATH45.9%—
LiveBench Math75.8%—
MATH Level 582.2%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Mistral Large 4 leads

Gemini 2.0 Flash (Feb 2025): 32.0 (#213), Mistral Large 4: 36.6 (#166)

Knowledge benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large 4
LMArena Expert13391447
GPQA Diamond64.1%—
Humanity's Last Exam6.6%—
SimpleQA Verified—20%
MMLU-Pro73.7%—
Confabulations12.4%—
GPQA (HELM)55.6%—
MMLU79.7%—

Multimodal Not comparable

Gemini 2.0 Flash (Feb 2025): 36.5 (#79), Mistral Large 4: —

Multimodal benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large 4
LMArena Vision1158—
GeoBench77%—

Multilingual Mistral Large 4 leads

Gemini 2.0 Flash (Feb 2025): 47.4 (#149), Mistral Large 4: 52.6 (#82)

Multilingual benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large 4
LMArena Non-English13421415
LMArena Chinese13731491
LMArena Russian13511414
LMArena French1391—
LMArena German1353—
LMArena Japanese1294—
LMArena Korean1313—
LMArena Spanish1363—

Instruction Following Too close to call

Gemini 2.0 Flash (Feb 2025): 74.4 (#97), Mistral Large 4: 75.0 (#76)

Instruction Following benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large 4
LMArena Instruction Following13361424
LiveBench Instruction Following85.8%—
IFEval84.1%—

Long Context Mistral Large 4 leads

Gemini 2.0 Flash (Feb 2025): 38.1 (#203), Mistral Large 4: 43.6 (#89)

Long Context benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large 4
LMArena Longer Query13441429
Fiction.LiveBench61.1%—

Writing & Preference Mistral Large 4 leads

Gemini 2.0 Flash (Feb 2025): 49.5 (#190), Mistral Large 4: 60.4 (#97)

Writing & Preference benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Mistral Large 4
LMArena Text13541427
LMArena Creative Writing13401361
LMArena Multi-Turn13501424
Short-Story Creative Writing73.8%—
EQ-Bench Creative Writing1128—
WildBench80%—
LiveBench Language51.3%—

Frequently asked questions

Is Gemini 2.0 Flash (Feb 2025) better than Mistral Large 4?

Mistral Large 4 is the stronger model overall, scoring 43.1 to 35.1 on the Noometry Index.

Is Gemini 2.0 Flash (Feb 2025) or Mistral Large 4 better for coding?

Mistral Large 4 scores higher on coding benchmarks: 48.6 versus 28.4 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Flash (Feb 2025) and Mistral Large 4 share?

12 benchmarks have published results for both models. Gemini 2.0 Flash (Feb 2025) has 54 scored results on Noometry and Mistral Large 4 has 15.

Related comparisons

Go deeper