Model comparison

Grok 4 Fast vs Mistral Small 3.1

Grok 4 Fast is the stronger model overall, scoring 39.4 to 31.7 on the Noometry Index.

Last verified . 18 shared benchmarks.

Grok 4 Fast xAI

39.4

Rank #167 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok 4 Fast scores higher in 7 categories and Mistral Small 3.1 in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4 Fast leads 38.9 to 14.7.
  • Mistral Small 3.1 has downloadable open weights; the other is API-only.

Side by side

Grok 4 Fast and Mistral Small 3.1 specifications
Grok 4 FastMistral Small 3.1
ProviderxAIMistral AI
Noometry Index39.431.7
Released2025-09-192025-03-17
WeightsProprietaryOpen
Context window—128K
Max output—102K
Input $ / M tokens—$0.35
Output $ / M tokens—$0.56
Results tracked3028

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3.1 leads

Grok 4 Fast: 32.5 (#271), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkGrok 4 FastMistral Small 3.1
LMArena Coding14291309
LMArena WebDev1159—
WeirdML42.9%—

Agentic & Tool Use Not comparable

Grok 4 Fast: 29.5 (#86), Mistral Small 3.1: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4 FastMistral Small 3.1
τ²-bench Banking15.7%—
Cybench30%—
LMArena Search1171—

Reasoning Grok 4 Fast leads

Grok 4 Fast: 22.2 (#201), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkGrok 4 FastMistral Small 3.1
LMArena Hard Prompts14121278
Epoch Capabilities Index144.2127.48
ARC-AGI-25.3%—
Kagi LLM Benchmark66.1%—
ARC-AGI-148.5%—
Chess Puzzles—1%
DTBench82.7%—
ForecastBench60.5—

Math Grok 4 Fast leads

Grok 4 Fast: 38.9 (#123), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkGrok 4 FastMistral Small 3.1
LMArena Math14191262
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%

Knowledge Grok 4 Fast leads

Grok 4 Fast: 32.0 (#214), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkGrok 4 FastMistral Small 3.1
LMArena Expert14111257
GPQA Diamond—41.9%
MMLU-Pro—61%
Vectara Hallucination Rate19.7%—
GPQA (HELM)—39.2%

Multimodal Not comparable

Grok 4 Fast: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkGrok 4 FastMistral Small 3.1
LMArena Vision—1136

Multilingual Grok 4 Fast leads

Grok 4 Fast: 51.3 (#111), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkGrok 4 FastMistral Small 3.1
LMArena Non-English13961255
LMArena Chinese14571253
LMArena French14311273
LMArena German13831266
LMArena Japanese13521208
LMArena Korean13571206
LMArena Russian13891263
LMArena Spanish14151283

Instruction Following Grok 4 Fast leads

Grok 4 Fast: 73.2 (#121), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkGrok 4 FastMistral Small 3.1
LMArena Instruction Following13871264
IFEval—75%

Long Context Grok 4 Fast leads

Grok 4 Fast: 63.2 (#3), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkGrok 4 FastMistral Small 3.1
LMArena Longer Query14151299
Fiction.LiveBench94.4%—

Writing & Preference Grok 4 Fast leads

Grok 4 Fast: 60.0 (#102), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkGrok 4 FastMistral Small 3.1
LMArena Text14071277
LMArena Creative Writing13871253
LMArena Multi-Turn14141270
EQ-Bench Creative Writing—761
WildBench—78.8%

Frequently asked questions

Is Grok 4 Fast better than Mistral Small 3.1?

Grok 4 Fast is the stronger model overall, scoring 39.4 to 31.7 on the Noometry Index.

Is Grok 4 Fast or Mistral Small 3.1 better for coding?

Mistral Small 3.1 scores higher on coding benchmarks: 38.3 versus 32.5 in the Noometry coding category.

How many benchmarks do Grok 4 Fast and Mistral Small 3.1 share?

18 benchmarks have published results for both models. Grok 4 Fast has 30 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper