Model comparison

Grok 2 Mini 2024 08 13 vs Mistral Large 3

Mistral Large 3 is the stronger model overall, scoring 39.1 to 37.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok 2 Mini 2024 08 13 xAI

37.7

Rank #198 Confirmed

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok 2 Mini 2024 08 13 scores higher in 2 categories and Mistral Large 3 in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Large 3 leads 60.0 to 47.3.
  • Mistral Large 3 has downloadable open weights; the other is API-only.

Side by side

Grok 2 Mini 2024 08 13 and Mistral Large 3 specifications
Grok 2 Mini 2024 08 13Mistral Large 3
ProviderxAIMistral AI
Noometry Index37.739.1
Released2024-08-132025-12-02
WeightsProprietaryOpen
Context window—262K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.75
Results tracked1724

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 37.0 (#199), Mistral Large 3: 34.4 (#237)

Coding benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large 3
LMArena Coding12691448
LMArena WebDev—1230

Reasoning Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 24.8 (#159), Mistral Large 3: 15.2 (#319)

Reasoning benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large 3
LMArena Hard Prompts12551429
Kagi LLM Benchmark—50.9%
NYT Connections (extended)—7.5%
Thematic Generalization—23%

Math Mistral Large 3 leads

Grok 2 Mini 2024 08 13: 35.4 (#185), Mistral Large 3: 38.7 (#129)

Math benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large 3
LMArena Math12651414

Knowledge Mistral Large 3 leads

Grok 2 Mini 2024 08 13: 33.9 (#200), Mistral Large 3: 36.0 (#177)

Knowledge benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large 3
LMArena Expert12381421
Vectara Hallucination Rate—14.5%

Multimodal Not comparable

Grok 2 Mini 2024 08 13: —, Mistral Large 3: 38.2 (#66)

Multimodal benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large 3
LMArena Vision—1221

Multilingual Mistral Large 3 leads

Grok 2 Mini 2024 08 13: 41.4 (#208), Mistral Large 3: 52.5 (#84)

Multilingual benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large 3
LMArena Non-English12571413
LMArena Chinese12621447
LMArena French12861455
LMArena German12731437
LMArena Japanese12131394
LMArena Korean11951384
LMArena Russian12611411
LMArena Spanish12791440

Instruction Following Mistral Large 3 leads

Grok 2 Mini 2024 08 13: 65.5 (#220), Mistral Large 3: 74.0 (#108)

Instruction Following benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large 3
LMArena Instruction Following12451403

Long Context Mistral Large 3 leads

Grok 2 Mini 2024 08 13: 38.4 (#196), Mistral Large 3: 43.1 (#105)

Long Context benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large 3
LMArena Longer Query12661413

Writing & Preference Mistral Large 3 leads

Grok 2 Mini 2024 08 13: 47.3 (#211), Mistral Large 3: 60.0 (#101)

Writing & Preference benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large 3
LMArena Text12811428
LMArena Creative Writing12421386
LMArena Multi-Turn12651429
EQ-Bench Creative Writing—1412

Frequently asked questions

Is Grok 2 Mini 2024 08 13 better than Mistral Large 3?

Mistral Large 3 is the stronger model overall, scoring 39.1 to 37.7 on the Noometry Index.

Is Grok 2 Mini 2024 08 13 or Mistral Large 3 better for coding?

Grok 2 Mini 2024 08 13 scores higher on coding benchmarks: 37.0 versus 34.4 in the Noometry coding category.

How many benchmarks do Grok 2 Mini 2024 08 13 and Mistral Large 3 share?

17 benchmarks have published results for both models. Grok 2 Mini 2024 08 13 has 17 scored results on Noometry and Mistral Large 3 has 24.

Related comparisons

Go deeper