Model comparison

Grok 4.1 vs Mistral Medium

Grok 4.1 is the stronger model overall, scoring 41.5 to 36.3 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok 4.1 scores higher in 8 categories and Mistral Medium in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.1 leads 39.5 to 25.0.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 and Mistral Medium specifications
Grok 4.1Mistral Medium
ProviderxAIMistral AI
Noometry Index41.536.3
Released2025-11-172023-12-11
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked1936

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.1: 33.7 (#253), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkGrok 4.1Mistral Medium
LMArena Coding14451434
FrontierCode—8%
LMArena WebDev1214—
SciCode—40.2%
WeirdML—43.7%
ALE-Bench—763.98

Agentic & Tool Use Grok 4.1 leads

Grok 4.1: 34.1 (#49), Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Mistral Medium
Berkeley Function Calling Leaderboard—37.7%
Cybench39%—

Reasoning Grok 4.1 leads

Grok 4.1: 29.5 (#91), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkGrok 4.1Mistral Medium
LMArena Hard Prompts14351426
Kagi LLM Benchmark—50%
CritPt—0%
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Grok 4.1 leads

Grok 4.1: 38.9 (#120), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkGrok 4.1Mistral Medium
LMArena Math14221408
OTIS Mock AIME 2024-2025—32.2%
ProofBench—9%
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Grok 4.1 leads

Grok 4.1: 39.5 (#133), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkGrok 4.1Mistral Medium
LMArena Expert14171408
GPQA Diamond—59.5%
Humanity's Last Exam—4.5%
Vectara Hallucination Rate—22.7%

Multimodal Not comparable

Grok 4.1: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkGrok 4.1Mistral Medium
LMArena Vision—1172

Multilingual Grok 4.1 leads

Grok 4.1: 53.4 (#68), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkGrok 4.1Mistral Medium
LMArena Non-English14251408
LMArena Chinese14651447
LMArena French14481459
LMArena German14461432
LMArena Japanese13971378
LMArena Korean14071380
LMArena Russian14341411
LMArena Spanish14381433

Instruction Following Too close to call

Grok 4.1: 73.8 (#111), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkGrok 4.1Mistral Medium
LMArena Instruction Following14001398

Long Context Too close to call

Grok 4.1: 43.2 (#100), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkGrok 4.1Mistral Medium
LMArena Longer Query14161406

Writing & Preference Grok 4.1 leads

Grok 4.1: 62.4 (#75), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkGrok 4.1Mistral Medium
LMArena Text14371424
LMArena Creative Writing14111391
LMArena Multi-Turn14371418
Short-Story Creative Writing—77.3%

Frequently asked questions

Is Grok 4.1 better than Mistral Medium?

Grok 4.1 is the stronger model overall, scoring 41.5 to 36.3 on the Noometry Index.

Is Grok 4.1 or Mistral Medium better for coding?

They score almost the same on coding (33.7 vs 34.2); test both on your own repository before choosing.

How many benchmarks do Grok 4.1 and Mistral Medium share?

17 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper