Model comparison

Grok Build 0.1 vs Mistral Nemo

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 26.4 on the Noometry Index. Mistral Nemo costs 8.3× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 20.7.
  • Mistral Nemo is cheaper at $0.15 / $0.15 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Grok Build 0.1 accepts more context: 256K tokens versus 128K.
  • Mistral Nemo has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Mistral Nemo specifications
Grok Build 0.1Mistral Nemo
ProviderxAIMistral AI
Noometry Index36.426.4
Released2026-04-162024-07-01
WeightsProprietaryOpen
Context window256K128K
Max output256K128K
Input $ / M tokens$1$0.15
Output $ / M tokens$2$0.15
Results tracked310

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Grok Build 0.1: 43.1 (#91), Mistral Nemo: —

Coding benchmarks
BenchmarkGrok Build 0.1Mistral Nemo
SciCode50.2%—

Agentic & Tool Use Too close to call

Grok Build 0.1: 22.7 (#129), Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Mistral Nemo
Berkeley Function Calling Leaderboard—27.6%
BALROG—17.6%
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkGrok Build 0.1Mistral Nemo
CritPt9.1%—
DTBench—48.6%
Epoch Capabilities Index—118.68
PIQA—83.5%

Math Not comparable

Grok Build 0.1: —, Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkGrok Build 0.1Mistral Nemo
MATH Level 5—10.8%
GSM8K—84.2%

Knowledge Not comparable

Grok Build 0.1: —, Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkGrok Build 0.1Mistral Nemo
GPQA Diamond—29.9%
BoolQ—82.5%

Writing & Preference Not comparable

Grok Build 0.1: —, Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Mistral Nemo
EQ-Bench Creative Writing—881

Frequently asked questions

Is Grok Build 0.1 better than Mistral Nemo?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 26.4 on the Noometry Index. Mistral Nemo costs 8.3× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Which is cheaper, Grok Build 0.1 or Mistral Nemo?

Mistral Nemo is cheaper. It lists at $0.15 per million input tokens and $0.15 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 128K.

How many benchmarks do Grok Build 0.1 and Mistral Nemo share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper