Model comparison

Grok Build 0.1 vs Mistral Small

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 33.4 on the Noometry Index. Mistral Small costs 4.8× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Last verified . 2 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Grok Build 0.1 scores higher in 2 categories and Mistral Small in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 19.8.
  • The biggest single-benchmark swing is SciCode: 50.2% for Grok Build 0.1 and 26.5% for Mistral Small.
  • Mistral Small is cheaper at $0.15 / $0.60 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Mistral Small accepts more context: 262K tokens versus 256K.
  • Mistral Small has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Mistral Small specifications
Grok Build 0.1Mistral Small
ProviderxAIMistral AI
Noometry Index36.433.4
Released2026-04-162024-02-26
WeightsProprietaryOpen
Context window256K262K
Max output256K256K
Input $ / M tokens$1$0.15
Output $ / M tokens$2$0.60
Results tracked339

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkGrok Build 0.1Mistral Small
SciCode50.2%26.5%
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
LMArena Coding—1362
BigCodeBench Complete—46.6%
ALE-Bench—497.62

Agentic & Tool Use Mistral Small leads

Grok Build 0.1: 22.7 (#129), Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Mistral Small
Berkeley Function Calling Leaderboard—37.1%
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkGrok Build 0.1Mistral Small
CritPt9.1%0%
Kagi LLM Benchmark—37.8%
LiveBench Reasoning—44.8%
LMArena Hard Prompts—1335
DTBench—70.9%
LiveBench Data Analysis—53.7%
LMCA—20.6%
LiveBench—44%

Math Not comparable

Grok Build 0.1: —, Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkGrok Build 0.1Mistral Small
OTIS Mock AIME 2024-2025—5.8%
LiveBench Math—39.9%
LMArena Math—1341
MATH Level 5—46.8%

Knowledge Not comparable

Grok Build 0.1: —, Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkGrok Build 0.1Mistral Small
GPQA Diamond—47.5%
Vectara Hallucination Rate—5.1%
LMArena Expert—1291
MMLU—68.7%

Multimodal Not comparable

Grok Build 0.1: —, Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkGrok Build 0.1Mistral Small
LMArena Vision—1142

Multilingual Not comparable

Grok Build 0.1: —, Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkGrok Build 0.1Mistral Small
LMArena Non-English—1315
LMArena Chinese—1340
LMArena French—1337
LMArena German—1340
LMArena Japanese—1275
LMArena Korean—1259
LMArena Russian—1324
LMArena Spanish—1346

Instruction Following Not comparable

Grok Build 0.1: —, Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Mistral Small
LiveBench Instruction Following—63.7%
LMArena Instruction Following—1310

Long Context Not comparable

Grok Build 0.1: —, Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkGrok Build 0.1Mistral Small
LMArena Longer Query—1327

Writing & Preference Not comparable

Grok Build 0.1: —, Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Mistral Small
LMArena Text—1338
LMArena Creative Writing—1305
LMArena Multi-Turn—1344
LiveBench Language—30.5%

Frequently asked questions

Is Grok Build 0.1 better than Mistral Small?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 33.4 on the Noometry Index. Mistral Small costs 4.8× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Which is cheaper, Grok Build 0.1 or Mistral Small?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or Mistral Small better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 34.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 256K.

How many benchmarks do Grok Build 0.1 and Mistral Small share?

2 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper