Model comparison

Grok 2 Mini 2024 08 13 vs Mistral Large

Grok 2 Mini 2024 08 13 is the stronger model overall, scoring 37.7 to 31.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok 2 Mini 2024 08 13 xAI

37.7

Rank #198 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok 2 Mini 2024 08 13 scores higher in 7 categories and Mistral Large in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 2 Mini 2024 08 13 leads 35.4 to 18.2.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Grok 2 Mini 2024 08 13 and Mistral Large specifications
Grok 2 Mini 2024 08 13Mistral Large
ProviderxAIMistral AI
Noometry Index37.731.9
Released2024-08-132024-02-26
WeightsProprietaryOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1751

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 37.0 (#199), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large
LMArena Coding12691277
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Grok 2 Mini 2024 08 13: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 24.8 (#159), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large
LMArena Hard Prompts12551257
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 35.4 (#185), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large
LMArena Math12651262
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 33.9 (#200), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large
LMArena Expert12381232
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
MMLU—80%

Multilingual Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 41.4 (#208), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large
LMArena Non-English12571237
LMArena Chinese12621240
LMArena French12861325
LMArena German12731254
LMArena Japanese12131188
LMArena Korean11951202
LMArena Russian12611257
LMArena Spanish12791268

Instruction Following Mistral Large leads

Grok 2 Mini 2024 08 13: 65.5 (#220), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large
LMArena Instruction Following12451249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Too close to call

Grok 2 Mini 2024 08 13: 38.4 (#196), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large
LMArena Longer Query12661261

Writing & Preference Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 47.3 (#211), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGrok 2 Mini 2024 08 13Mistral Large
LMArena Text12811266
LMArena Creative Writing12421243
LMArena Multi-Turn12651260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Grok 2 Mini 2024 08 13 better than Mistral Large?

Grok 2 Mini 2024 08 13 is the stronger model overall, scoring 37.7 to 31.9 on the Noometry Index.

Is Grok 2 Mini 2024 08 13 or Mistral Large better for coding?

Grok 2 Mini 2024 08 13 scores higher on coding benchmarks: 37.0 versus 34.3 in the Noometry coding category.

How many benchmarks do Grok 2 Mini 2024 08 13 and Mistral Large share?

17 benchmarks have published results for both models. Grok 2 Mini 2024 08 13 has 17 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper