Model comparison

Grok 4.1 vs Mistral Large 3

Grok 4.1 is the stronger model overall, scoring 41.5 to 39.1 on the Noometry Index.

Last verified . 18 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok 4.1 scores higher in 6 categories and Mistral Large 3 in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 leads 29.5 to 15.2.
  • Mistral Large 3 has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 and Mistral Large 3 specifications
Grok 4.1Mistral Large 3
ProviderxAIMistral AI
Noometry Index41.539.1
Released2025-11-172025-12-02
WeightsProprietaryOpen
Context window—262K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.75
Results tracked1924

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.1: 33.7 (#253), Mistral Large 3: 34.4 (#237)

Coding benchmarks
BenchmarkGrok 4.1Mistral Large 3
LMArena WebDev12141230
LMArena Coding14451448

Agentic & Tool Use Not comparable

Grok 4.1: 34.1 (#49), Mistral Large 3: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Mistral Large 3
Cybench39%—

Reasoning Grok 4.1 leads

Grok 4.1: 29.5 (#91), Mistral Large 3: 15.2 (#319)

Reasoning benchmarks
BenchmarkGrok 4.1Mistral Large 3
LMArena Hard Prompts14351429
Kagi LLM Benchmark—50.9%
NYT Connections (extended)—7.5%
Thematic Generalization—23%

Math Too close to call

Grok 4.1: 38.9 (#120), Mistral Large 3: 38.7 (#129)

Math benchmarks
BenchmarkGrok 4.1Mistral Large 3
LMArena Math14221414

Knowledge Grok 4.1 leads

Grok 4.1: 39.5 (#133), Mistral Large 3: 36.0 (#177)

Knowledge benchmarks
BenchmarkGrok 4.1Mistral Large 3
LMArena Expert14171421
Vectara Hallucination Rate—14.5%

Multimodal Not comparable

Grok 4.1: —, Mistral Large 3: 38.2 (#66)

Multimodal benchmarks
BenchmarkGrok 4.1Mistral Large 3
LMArena Vision—1221

Multilingual Too close to call

Grok 4.1: 53.4 (#68), Mistral Large 3: 52.5 (#84)

Multilingual benchmarks
BenchmarkGrok 4.1Mistral Large 3
LMArena Non-English14251413
LMArena Chinese14651447
LMArena French14481455
LMArena German14461437
LMArena Japanese13971394
LMArena Korean14071384
LMArena Russian14341411
LMArena Spanish14381440

Instruction Following Too close to call

Grok 4.1: 73.8 (#111), Mistral Large 3: 74.0 (#108)

Instruction Following benchmarks
BenchmarkGrok 4.1Mistral Large 3
LMArena Instruction Following14001403

Long Context Too close to call

Grok 4.1: 43.2 (#100), Mistral Large 3: 43.1 (#105)

Long Context benchmarks
BenchmarkGrok 4.1Mistral Large 3
LMArena Longer Query14161413

Writing & Preference Grok 4.1 leads

Grok 4.1: 62.4 (#75), Mistral Large 3: 60.0 (#101)

Writing & Preference benchmarks
BenchmarkGrok 4.1Mistral Large 3
LMArena Text14371428
LMArena Creative Writing14111386
LMArena Multi-Turn14371429
EQ-Bench Creative Writing—1412

Frequently asked questions

Is Grok 4.1 better than Mistral Large 3?

Grok 4.1 is the stronger model overall, scoring 41.5 to 39.1 on the Noometry Index.

Is Grok 4.1 or Mistral Large 3 better for coding?

They score almost the same on coding (33.7 vs 34.4); test both on your own repository before choosing.

How many benchmarks do Grok 4.1 and Mistral Large 3 share?

18 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Mistral Large 3 has 24.

Related comparisons

Go deeper