Model comparison

Grok 4.1 Fast vs Mistral Small 3.1

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 31.7 on the Noometry Index.

Last verified . 19 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Grok 4.1 Fast scores higher in 8 categories and Mistral Small 3.1 in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 19.7.
  • Grok 4.1 Fast is cheaper at $0.20 / $0.50 per million input/output tokens, against $0.35 / $0.56 for Mistral Small 3.1.
  • Mistral Small 3.1 has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 Fast and Mistral Small 3.1 specifications
Grok 4.1 FastMistral Small 3.1
ProviderxAIMistral AI
Noometry Index41.431.7
Released2025-06-272025-03-17
WeightsProprietaryOpen
Context window128K128K
Max output30K102K
Input $ / M tokens$0.20$0.35
Output $ / M tokens$0.50$0.56
Results tracked3228

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3.1 leads

Grok 4.1 Fast: 34.1 (#245), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkGrok 4.1 FastMistral Small 3.1
LMArena Coding14111309
LMArena WebDev1242—
ALE-Bench394.93—

Agentic & Tool Use Not comparable

Grok 4.1 Fast: 36.3 (#39), Mistral Small 3.1: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1 FastMistral Small 3.1
Berkeley Function Calling Leaderboard69.6%—
τ²-bench Banking13.1%—
LMArena Search1171—
Vending-Bench 21,107—

Reasoning Grok 4.1 Fast leads

Grok 4.1 Fast: 43.4 (#49), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkGrok 4.1 FastMistral Small 3.1
LMArena Hard Prompts14071278
SimpleBench56%—
NYT Connections (extended)87.4%—
Chess Puzzles—1%
DTBench87.7%—
Epoch Capabilities Index—127.48
ForecastBench61—

Math Grok 4.1 Fast leads

Grok 4.1 Fast: 31.9 (#221), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkGrok 4.1 FastMistral Small 3.1
LMArena Math14081262
MathArena Final-Answer Competitions60.9%—
OTIS Mock AIME 2024-2025—3.9%
ProofBench4%—
Omni-MATH—24.8%

Knowledge Grok 4.1 Fast leads

Grok 4.1 Fast: 33.1 (#207), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkGrok 4.1 FastMistral Small 3.1
LMArena Expert13991257
GPQA Diamond—41.9%
MMLU-Pro—61%
Vectara Hallucination Rate17.8%—
GPQA (HELM)—39.2%

Multimodal Grok 4.1 Fast leads

Grok 4.1 Fast: 37.0 (#76), Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkGrok 4.1 FastMistral Small 3.1
LMArena Vision12011136

Multilingual Grok 4.1 Fast leads

Grok 4.1 Fast: 51.0 (#114), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkGrok 4.1 FastMistral Small 3.1
LMArena Non-English13911255
LMArena Chinese14411253
LMArena French14151273
LMArena German14041266
LMArena Japanese13491208
LMArena Korean13611206
LMArena Russian13871263
LMArena Spanish14131283

Instruction Following Grok 4.1 Fast leads

Grok 4.1 Fast: 72.7 (#133), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkGrok 4.1 FastMistral Small 3.1
LMArena Instruction Following13761264
IFEval—75%

Long Context Grok 4.1 Fast leads

Grok 4.1 Fast: 42.4 (#126), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkGrok 4.1 FastMistral Small 3.1
LMArena Longer Query13901299

Writing & Preference Grok 4.1 Fast leads

Grok 4.1 Fast: 57.2 (#131), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkGrok 4.1 FastMistral Small 3.1
LMArena Text14081277
LMArena Creative Writing13941253
EQ-Bench Creative Writing1327761
LMArena Multi-Turn13891270
WildBench—78.8%

Frequently asked questions

Is Grok 4.1 Fast better than Mistral Small 3.1?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 31.7 on the Noometry Index.

Which is cheaper, Grok 4.1 Fast or Mistral Small 3.1?

Grok 4.1 Fast is cheaper. It lists at $0.20 per million input tokens and $0.50 per million output tokens; Mistral Small 3.1 lists at $0.35 and $0.56.

Is Grok 4.1 Fast or Mistral Small 3.1 better for coding?

Mistral Small 3.1 scores higher on coding benchmarks: 38.3 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Grok 4.1 Fast and Mistral Small 3.1 share?

19 benchmarks have published results for both models. Grok 4.1 Fast has 32 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper