Model comparison

Grok 4.1 vs Mistral Large

Grok 4.1 is the stronger model overall, scoring 41.5 to 31.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok 4.1 scores higher in 8 categories and Mistral Large in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Grok 4.1 leads 62.4 to 40.7.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 and Mistral Large specifications
Grok 4.1Mistral Large
ProviderxAIMistral AI
Noometry Index41.531.9
Released2025-11-172024-02-26
WeightsProprietaryOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1951

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.1: 33.7 (#253), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGrok 4.1Mistral Large
LMArena Coding14451277
LMArena WebDev1214—
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Grok 4.1 leads

Grok 4.1: 34.1 (#49), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Mistral Large
Berkeley Function Calling Leaderboard—38.4%
Cybench39%—

Reasoning Grok 4.1 leads

Grok 4.1: 29.5 (#91), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGrok 4.1Mistral Large
LMArena Hard Prompts14351257
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Grok 4.1 leads

Grok 4.1: 38.9 (#120), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGrok 4.1Mistral Large
LMArena Math14221262
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Grok 4.1 leads

Grok 4.1: 39.5 (#133), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGrok 4.1Mistral Large
LMArena Expert14171232
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
MMLU—80%

Multilingual Grok 4.1 leads

Grok 4.1: 53.4 (#68), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGrok 4.1Mistral Large
LMArena Non-English14251237
LMArena Chinese14651240
LMArena French14481325
LMArena German14461254
LMArena Japanese13971188
LMArena Korean14071202
LMArena Russian14341257
LMArena Spanish14381268

Instruction Following Grok 4.1 leads

Grok 4.1: 73.8 (#111), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGrok 4.1Mistral Large
LMArena Instruction Following14001249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Grok 4.1 leads

Grok 4.1: 43.2 (#100), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGrok 4.1Mistral Large
LMArena Longer Query14161261

Writing & Preference Grok 4.1 leads

Grok 4.1: 62.4 (#75), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGrok 4.1Mistral Large
LMArena Text14371266
LMArena Creative Writing14111243
LMArena Multi-Turn14371260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Grok 4.1 better than Mistral Large?

Grok 4.1 is the stronger model overall, scoring 41.5 to 31.9 on the Noometry Index.

Is Grok 4.1 or Mistral Large better for coding?

They score almost the same on coding (33.7 vs 34.3); test both on your own repository before choosing.

How many benchmarks do Grok 4.1 and Mistral Large share?

17 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper