Model comparison

Grok Build 0.1 vs Mistral Large

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 31.9 on the Noometry Index.

Last verified . 2 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Grok Build 0.1 scores higher in 2 categories and Mistral Large in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 15.8.
  • The biggest single-benchmark swing is SciCode: 50.2% for Grok Build 0.1 and 36.2% for Mistral Large.
  • Grok Build 0.1 is cheaper at $1 / $2 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Grok Build 0.1 accepts more context: 256K tokens versus 131K.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Mistral Large specifications
Grok Build 0.1Mistral Large
ProviderxAIMistral AI
Noometry Index36.431.9
Released2026-04-162024-02-26
WeightsProprietaryOpen
Context window256K131K
Max output256K16K
Input $ / M tokens$1$2
Output $ / M tokens$2$6
Results tracked351

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGrok Build 0.1Mistral Large
SciCode50.2%36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
LMArena Coding—1277
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Mistral Large leads

Grok Build 0.1: 22.7 (#129), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Mistral Large
Berkeley Function Calling Leaderboard—38.4%
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGrok Build 0.1Mistral Large
CritPt9.1%0%
SimpleBench—22.5%
LiveBench Reasoning—43.5%
LMArena Hard Prompts—1257
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Not comparable

Grok Build 0.1: —, Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGrok Build 0.1Mistral Large
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
LMArena Math—1262
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Not comparable

Grok Build 0.1: —, Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGrok Build 0.1Mistral Large
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
LMArena Expert—1232
MMLU—80%

Multilingual Not comparable

Grok Build 0.1: —, Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGrok Build 0.1Mistral Large
LMArena Non-English—1237
LMArena Chinese—1240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Not comparable

Grok Build 0.1: —, Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Mistral Large
LiveBench Instruction Following—67.9%
IFEval—87.7%
LMArena Instruction Following—1249

Long Context Not comparable

Grok Build 0.1: —, Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGrok Build 0.1Mistral Large
LMArena Longer Query—1261

Writing & Preference Not comparable

Grok Build 0.1: —, Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Mistral Large
LMArena Text—1266
LMArena Creative Writing—1243
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LMArena Multi-Turn—1260
LiveBench Language—39.4%

Frequently asked questions

Is Grok Build 0.1 better than Mistral Large?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 31.9 on the Noometry Index.

Which is cheaper, Grok Build 0.1 or Mistral Large?

Grok Build 0.1 is cheaper. It lists at $1 per million input tokens and $2 per million output tokens; Mistral Large lists at $2 and $6.

Is Grok Build 0.1 or Mistral Large better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 131K.

How many benchmarks do Grok Build 0.1 and Mistral Large share?

2 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper