Model comparison

Grok Build 0.1 vs Mistral

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 29.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Mistral Mistral AI

29.9

Rank #303 Confirmed

Summary

  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 22.2.

Side by side

Grok Build 0.1 and Mistral specifications
Grok Build 0.1Mistral
ProviderxAIMistral AI
Noometry Index36.429.9
Released2026-04-16—
WeightsProprietaryProprietary
Context window256K—
Max output256K—
Input $ / M tokens$1—
Output $ / M tokens$2—
Results tracked322

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Mistral: 33.8 (#250)

Coding benchmarks
BenchmarkGrok Build 0.1Mistral
SciCode50.2%—
LMArena Coding—1162

Agentic & Tool Use Not comparable

Grok Build 0.1: 22.7 (#129), Mistral: —

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Mistral
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Mistral: 22.2 (#200)

Reasoning benchmarks
BenchmarkGrok Build 0.1Mistral
CritPt9.1%—
LMArena Hard Prompts—1149

Math Not comparable

Grok Build 0.1: —, Mistral: 22.3 (#278)

Math benchmarks
BenchmarkGrok Build 0.1Mistral
Omni-MATH—7.2%
LMArena Math—1180

Knowledge Not comparable

Grok Build 0.1: —, Mistral: 16.6 (#288)

Knowledge benchmarks
BenchmarkGrok Build 0.1Mistral
MMLU-Pro—27.7%
GPQA (HELM)—30.3%
LMArena Expert—1125

Multilingual Not comparable

Grok Build 0.1: —, Mistral: 32.8 (#254)

Multilingual benchmarks
BenchmarkGrok Build 0.1Mistral
LMArena Non-English—1129
LMArena Chinese—1109
LMArena French—1180
LMArena German—1155
LMArena Japanese—1013
LMArena Korean—1032
LMArena Russian—1168
LMArena Spanish—1143

Instruction Following Not comparable

Grok Build 0.1: —, Mistral: 52.6 (#288)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Mistral
IFEval—56.8%
LMArena Instruction Following—1152

Long Context Not comparable

Grok Build 0.1: —, Mistral: 35.0 (#245)

Long Context benchmarks
BenchmarkGrok Build 0.1Mistral
LMArena Longer Query—1153

Writing & Preference Not comparable

Grok Build 0.1: —, Mistral: 37.0 (#260)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Mistral
LMArena Text—1165
LMArena Creative Writing—1158
WildBench—66%
LMArena Multi-Turn—1147

Frequently asked questions

Is Grok Build 0.1 better than Mistral?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 29.9 on the Noometry Index.

Is Grok Build 0.1 or Mistral better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 33.8 in the Noometry coding category.

How many benchmarks do Grok Build 0.1 and Mistral share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Mistral has 22.

Related comparisons

Go deeper