Model comparison

Gemini 3.1 Pro Preview vs Mistral Medium

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 36.3 on the Noometry Index. Mistral Medium costs 1.5× less per token, which makes it the better buy when Gemini 3.1 Pro Preview's lead doesn't matter for your workload.

Last verified . 30 shared benchmarks.

Gemini 3.1 Pro Preview Google

56.7

Rank #23 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Gemini 3.1 Pro Preview scores higher in 10 categories and Mistral Medium in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.1 Pro Preview leads 71.7 to 24.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 95.6% for Gemini 3.1 Pro Preview and 32.2% for Mistral Medium.
  • Mistral Medium is cheaper at $1.50 / $7.50 per million input/output tokens, against $2 / $12 for Gemini 3.1 Pro Preview.
  • Gemini 3.1 Pro Preview accepts more context: 1.05M tokens versus 262K.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Gemini 3.1 Pro Preview and Mistral Medium specifications
Gemini 3.1 Pro PreviewMistral Medium
ProviderGoogleMistral AI
Noometry Index56.736.3
Released2026-02-192023-12-11
WeightsProprietaryOpen
Context window1.05M262K
Max output66K262K
Input $ / M tokens$2$1.50
Output $ / M tokens$12$7.50
Results tracked7136

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 42.5 (#99), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Medium
SciCode58.9%40.2%
WeirdML72.1%43.7%
LMArena Coding14841434
ALE-Bench1,161763.98
SWE-bench Verified75.6%—
DeepSWE11.7%—
FrontierCode—8%
LMArena WebDev1447—
GSO22.6%—
MirrorCode8.9%—
AlgoTune2.02—

Agentic & Tool Use Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 37.7 (#34), Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Medium
Terminal-Bench80.2%—
APEX-Agents35.3%—
Berkeley Function Calling Leaderboard—37.7%
τ²-bench Banking26%—
DeepResearch Bench47.8%—
PostTrainBench22%—
BALROG57%—
ExploitBench26.1%—
GBAEval0.8%—
GDP.pdf17%—
LMArena Search1211—
METR Time Horizons77%—
Vending-Bench 23,774—

Reasoning Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 71.7 (#12), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Medium
CritPt17.7%0%
LMArena Hard Prompts14851426
DTBench97.1%75.5%
LMCA53.8%26.1%
ARC-AGI-277.1%—
SimpleBench79.6%—
Kagi LLM Benchmark—50%
NYT Connections (extended)97.4%—
ARC-AGI-198%—
Chess Puzzles55%—
EnigmaEval36.8%—
Thematic Generalization79.4%—
EBR-Bench14.3%—
Mystery Game Puzzles34%—
Surface Evolver Bench—26.9%
Epoch Capabilities Index154.77—
ForecastBench59—

Math Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 62.1 (#34), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Medium
OTIS Mock AIME 2024-202595.6%32.2%
ProofBench26%9%
LMArena Math14851408
FrontierMath (Feb 2025 set)36.9%0.3%
FrontierMath (Tiers 1-3)59.6%—
FrontierMath Tier 426.8%—
MathArena Final-Answer Competitions86.5%—
MATH Level 5—81.6%
FrontierMath Tier 4 (v1)16.7%—

Knowledge Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 71.8 (#3), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Medium
GPQA Diamond94.4%59.5%
Humanity's Last Exam46.4%4.5%
Vectara Hallucination Rate10.4%22.7%
LMArena Expert14851408
SimpleQA Verified73.5%—

Multimodal Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 37.9 (#69), Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Medium
LMArena Vision12961172
Blueprint-Bench 226.5%—
Furniture Assembly26.7%—
LMArena Document1444—

Multilingual Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 57.0 (#12), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Medium
LMArena Non-English14771408
LMArena Chinese15291447
LMArena French14871459
LMArena German14911432
LMArena Japanese14931378
LMArena Korean14551380
LMArena Russian14981411
LMArena Spanish14791433

Instruction Following Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 77.0 (#32), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Medium
LMArena Instruction Following14661398

Long Context Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 47.4 (#18), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Medium
LMArena Longer Query14831406
CL-bench20.8%—
CL-bench Life16.9%—

Writing & Preference Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 66.1 (#37), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Medium
LMArena Text14811424
LMArena Creative Writing14821391
LMArena Multi-Turn14881418
Short-Story Creative Writing—77.3%
EQ-Bench Creative Writing1491—
EQ-Bench 41142—

Frequently asked questions

Is Gemini 3.1 Pro Preview better than Mistral Medium?

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 36.3 on the Noometry Index. Mistral Medium costs 1.5× less per token, which makes it the better buy when Gemini 3.1 Pro Preview's lead doesn't matter for your workload.

Which is cheaper, Gemini 3.1 Pro Preview or Mistral Medium?

Mistral Medium is cheaper. It lists at $1.50 per million input tokens and $7.50 per million output tokens; Gemini 3.1 Pro Preview lists at $2 and $12.

Is Gemini 3.1 Pro Preview or Mistral Medium better for coding?

Gemini 3.1 Pro Preview scores higher on coding benchmarks: 42.5 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

Gemini 3.1 Pro Preview does, with 1.05M tokens against 262K.

How many benchmarks do Gemini 3.1 Pro Preview and Mistral Medium share?

30 benchmarks have published results for both models. Gemini 3.1 Pro Preview has 71 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper