Model comparison

Grok Build 0.1 vs Pixtral Large

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 32.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Summary

  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 21.7.
  • Grok Build 0.1 is cheaper at $1 / $2 per million input/output tokens, against $2 / $6 for Pixtral Large.
  • Grok Build 0.1 accepts more context: 256K tokens versus 128K.
  • Pixtral Large has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Pixtral Large specifications
Grok Build 0.1Pixtral Large
ProviderxAIMistral AI
Noometry Index36.432.2
Released2026-04-162024-11-01
WeightsProprietaryOpen
Context window256K128K
Max output256K128K
Input $ / M tokens$1$2
Output $ / M tokens$2$6
Results tracked33

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Grok Build 0.1: 43.1 (#91), Pixtral Large: —

Coding benchmarks
BenchmarkGrok Build 0.1Pixtral Large
SciCode50.2%—

Agentic & Tool Use Not comparable

Grok Build 0.1: 22.7 (#129), Pixtral Large: —

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Pixtral Large
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Pixtral Large: 21.7 (#218)

Reasoning benchmarks
BenchmarkGrok Build 0.1Pixtral Large
CritPt9.1%—
EnigmaEval—0.8%

Multimodal Not comparable

Grok Build 0.1: —, Pixtral Large: 30.6 (#111)

Multimodal benchmarks
BenchmarkGrok Build 0.1Pixtral Large
LMArena Vision—1089

Writing & Preference Not comparable

Grok Build 0.1: —, Pixtral Large: 32.9 (#278)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Pixtral Large
EQ-Bench Creative Writing—988

Frequently asked questions

Is Grok Build 0.1 better than Pixtral Large?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 32.2 on the Noometry Index.

Which is cheaper, Grok Build 0.1 or Pixtral Large?

Grok Build 0.1 is cheaper. It lists at $1 per million input tokens and $2 per million output tokens; Pixtral Large lists at $2 and $6.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 128K.

How many benchmarks do Grok Build 0.1 and Pixtral Large share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Pixtral Large has 3.

Related comparisons

Go deeper