Model comparison

Gemini 3.5 Flash vs Mistral Large 4

Gemini 3.5 Flash is the stronger model overall, scoring 54.2 to 43.1 on the Noometry Index. Mistral Large 4 costs 3.3× less per token, which makes it the better buy when Gemini 3.5 Flash's lead doesn't matter for your workload.

Last verified . 15 shared benchmarks.

Gemini 3.5 Flash Google

54.2

Rank #32 Confirmed

Mistral Large 4 Mistral AI

43.1

Rank #99 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Gemini 3.5 Flash scores higher in 8 categories and Mistral Large 4 in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.5 Flash leads 62.8 to 22.5.
  • The biggest single-benchmark swing is NYT Connections (extended): 92.6% for Gemini 3.5 Flash and 27.4% for Mistral Large 4.
  • Mistral Large 4 is cheaper at $0.68 / $2.09 per million input/output tokens, against $1.50 / $9 for Gemini 3.5 Flash.

Side by side

Gemini 3.5 Flash and Mistral Large 4 specifications
Gemini 3.5 FlashMistral Large 4
ProviderGoogleMistral AI
Noometry Index54.243.1
Released2026-05-192026-10-06
WeightsProprietaryProprietary
Context window1.05M1.05M
Max output66K262K
Input $ / M tokens$1.50$0.68
Output $ / M tokens$9$2.09
Results tracked5415

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemini 3.5 Flash: 49.4 (#49), Mistral Large 4: 48.6 (#57)

Coding benchmarks
BenchmarkGemini 3.5 FlashMistral Large 4
LMArena WebDev14991541
LMArena Coding14921475
SWE-bench Verified79.3%—
DeepSWE37.4%—
SciCode53.1%—
WeirdML62.6%—
ALE-Bench911.02—

Agentic & Tool Use Not comparable

Gemini 3.5 Flash: 24.7 (#114), Mistral Large 4: —

Agentic & Tool Use benchmarks
BenchmarkGemini 3.5 FlashMistral Large 4
APEX-Agents27.5%—
GBAEval6.7%—
GDP.pdf14%—
Vending-Bench 25,396—

Reasoning Gemini 3.5 Flash leads

Gemini 3.5 Flash: 62.8 (#18), Mistral Large 4: 22.5 (#192)

Reasoning benchmarks
BenchmarkGemini 3.5 FlashMistral Large 4
NYT Connections (extended)92.6%27.4%
LMArena Hard Prompts14881444
ARC-AGI-272.1%—
SimpleBench76.7%—
ARC-AGI-192.5%—
CritPt13.1%—
Chess Puzzles50%—
EnigmaEval25.4%—
EBR-Bench4.8%—
Mystery Game Puzzles32%—
DTBench94.7%—
LMCA47.1%—
Surface Evolver Bench58.1%—
Epoch Capabilities Index154.46—
ForecastBench59—

Math Gemini 3.5 Flash leads

Gemini 3.5 Flash: 60.7 (#36), Mistral Large 4: 40.4 (#91)

Knowledge Gemini 3.5 Flash leads

Gemini 3.5 Flash: 66.3 (#11), Mistral Large 4: 36.6 (#166)

Knowledge benchmarks
BenchmarkGemini 3.5 FlashMistral Large 4
SimpleQA Verified66.2%20%
LMArena Expert14951447
GPQA Diamond92.8%—

Multimodal Not comparable

Gemini 3.5 Flash: 45.7 (#15), Mistral Large 4: —

Multimodal benchmarks
BenchmarkGemini 3.5 FlashMistral Large 4
LMArena Vision1310—
Blueprint-Bench 233.6%—
LMArena Document1463—

Multilingual Gemini 3.5 Flash leads

Gemini 3.5 Flash: 57.0 (#13), Mistral Large 4: 52.6 (#82)

Multilingual benchmarks
BenchmarkGemini 3.5 FlashMistral Large 4
LMArena Non-English14761415
LMArena Chinese15261491
LMArena Russian14931414
LMArena French1490—
LMArena German1492—
LMArena Japanese1486—
LMArena Korean1451—
LMArena Spanish1480—

Instruction Following Gemini 3.5 Flash leads

Gemini 3.5 Flash: 77.0 (#30), Mistral Large 4: 75.0 (#76)

Instruction Following benchmarks
BenchmarkGemini 3.5 FlashMistral Large 4
LMArena Instruction Following14671424

Long Context Gemini 3.5 Flash leads

Gemini 3.5 Flash: 45.4 (#38), Mistral Large 4: 43.6 (#89)

Long Context benchmarks
BenchmarkGemini 3.5 FlashMistral Large 4
LMArena Longer Query14821429

Writing & Preference Gemini 3.5 Flash leads

Gemini 3.5 Flash: 65.5 (#47), Mistral Large 4: 60.4 (#97)

Writing & Preference benchmarks
BenchmarkGemini 3.5 FlashMistral Large 4
LMArena Text14821427
LMArena Creative Writing14701361
LMArena Multi-Turn14811424
EQ-Bench 41087—

Frequently asked questions

Is Gemini 3.5 Flash better than Mistral Large 4?

Gemini 3.5 Flash is the stronger model overall, scoring 54.2 to 43.1 on the Noometry Index. Mistral Large 4 costs 3.3× less per token, which makes it the better buy when Gemini 3.5 Flash's lead doesn't matter for your workload.

Which is cheaper, Gemini 3.5 Flash or Mistral Large 4?

Mistral Large 4 is cheaper. It lists at $0.68 per million input tokens and $2.09 per million output tokens; Gemini 3.5 Flash lists at $1.50 and $9.

Is Gemini 3.5 Flash or Mistral Large 4 better for coding?

They score almost the same on coding (49.4 vs 48.6); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 1.05M tokens.

How many benchmarks do Gemini 3.5 Flash and Mistral Large 4 share?

15 benchmarks have published results for both models. Gemini 3.5 Flash has 54 scored results on Noometry and Mistral Large 4 has 15.

Related comparisons

Go deeper