Model comparison

Grok 4.1 Fast vs Mistral Small

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 33.4 on the Noometry Index.

Last verified . 22 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Grok 4.1 Fast scores higher in 10 categories and Mistral Small in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 19.8.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 69.6% for Grok 4.1 Fast and 37.1% for Mistral Small.
  • Both cost about the same: $0.20 input and $0.50 output per million tokens.
  • Mistral Small accepts more context: 262K tokens versus 128K.
  • Mistral Small has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 Fast and Mistral Small specifications
Grok 4.1 FastMistral Small
ProviderxAIMistral AI
Noometry Index41.433.4
Released2025-06-272024-02-26
WeightsProprietaryOpen
Context window128K262K
Max output30K256K
Input $ / M tokens$0.20$0.15
Output $ / M tokens$0.50$0.60
Results tracked3239

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.1 Fast: 34.1 (#245), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkGrok 4.1 FastMistral Small
LMArena Coding14111362
ALE-Bench394.93497.62
LMArena WebDev1242—
SciCode—26.5%
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
BigCodeBench Complete—46.6%

Agentic & Tool Use Grok 4.1 Fast leads

Grok 4.1 Fast: 36.3 (#39), Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1 FastMistral Small
Berkeley Function Calling Leaderboard69.6%37.1%
τ²-bench Banking13.1%—
LMArena Search1171—
Vending-Bench 21,107—

Reasoning Grok 4.1 Fast leads

Grok 4.1 Fast: 43.4 (#49), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkGrok 4.1 FastMistral Small
LMArena Hard Prompts14071335
DTBench87.7%70.9%
SimpleBench56%—
Kagi LLM Benchmark—37.8%
NYT Connections (extended)87.4%—
CritPt—0%
LiveBench Reasoning—44.8%
LiveBench Data Analysis—53.7%
LMCA—20.6%
ForecastBench61—
LiveBench—44%

Math Grok 4.1 Fast leads

Grok 4.1 Fast: 31.9 (#221), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkGrok 4.1 FastMistral Small
LMArena Math14081341
MathArena Final-Answer Competitions60.9%—
OTIS Mock AIME 2024-2025—5.8%
ProofBench4%—
LiveBench Math—39.9%
MATH Level 5—46.8%

Knowledge Grok 4.1 Fast leads

Grok 4.1 Fast: 33.1 (#207), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkGrok 4.1 FastMistral Small
Vectara Hallucination Rate17.8%5.1%
LMArena Expert13991291
GPQA Diamond—47.5%
MMLU—68.7%

Multimodal Grok 4.1 Fast leads

Grok 4.1 Fast: 37.0 (#76), Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkGrok 4.1 FastMistral Small
LMArena Vision12011142

Multilingual Grok 4.1 Fast leads

Grok 4.1 Fast: 51.0 (#114), Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkGrok 4.1 FastMistral Small
LMArena Non-English13911315
LMArena Chinese14411340
LMArena French14151337
LMArena German14041340
LMArena Japanese13491275
LMArena Korean13611259
LMArena Russian13871324
LMArena Spanish14131346

Instruction Following Grok 4.1 Fast leads

Grok 4.1 Fast: 72.7 (#133), Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkGrok 4.1 FastMistral Small
LMArena Instruction Following13761310
LiveBench Instruction Following—63.7%

Long Context Grok 4.1 Fast leads

Grok 4.1 Fast: 42.4 (#126), Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkGrok 4.1 FastMistral Small
LMArena Longer Query13901327

Writing & Preference Grok 4.1 Fast leads

Grok 4.1 Fast: 57.2 (#131), Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkGrok 4.1 FastMistral Small
LMArena Text14081338
LMArena Creative Writing13941305
LMArena Multi-Turn13891344
EQ-Bench Creative Writing1327—
LiveBench Language—30.5%

Frequently asked questions

Is Grok 4.1 Fast better than Mistral Small?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 33.4 on the Noometry Index.

Which is cheaper, Grok 4.1 Fast or Mistral Small?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Grok 4.1 Fast lists at $0.20 and $0.50.

Is Grok 4.1 Fast or Mistral Small better for coding?

They score almost the same on coding (34.1 vs 34.0); test both on your own repository before choosing.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 128K.

How many benchmarks do Grok 4.1 Fast and Mistral Small share?

22 benchmarks have published results for both models. Grok 4.1 Fast has 32 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper