Model comparison

Grok Build 0.1 vs Mixtral 8x7B

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 27.1 on the Noometry Index. Mixtral 8x7B costs 1.8× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Summary

  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 18.2.
  • Mixtral 8x7B is cheaper at $0.70 / $0.70 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Grok Build 0.1 accepts more context: 256K tokens versus 32K.
  • Mixtral 8x7B has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Mixtral 8x7B specifications
Grok Build 0.1Mixtral 8x7B
ProviderxAIMistral AI
Noometry Index36.427.1
Released2026-04-162023-12-11
WeightsProprietaryOpen
Context window256K32K
Max output256K32K
Input $ / M tokens$1$0.70
Output $ / M tokens$2$0.70
Results tracked338

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Mixtral 8x7B: 32.8 (#269)

Coding benchmarks
BenchmarkGrok Build 0.1Mixtral 8x7B
SciCode50.2%—
LMArena Coding—1126
HumanEval+—39.6%
MBPP+—49.7%

Agentic & Tool Use Not comparable

Grok Build 0.1: 22.7 (#129), Mixtral 8x7B: —

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Mixtral 8x7B
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Mixtral 8x7B: 18.2 (#285)

Reasoning benchmarks
BenchmarkGrok Build 0.1Mixtral 8x7B
CritPt9.1%—
LMArena Hard Prompts—1115
DTBench—49.6%
Adversarial NLI—55.2%
Epoch Capabilities Index—118.47
ForecastBench—56.3
HellaSwag—86.7%
PIQA—83.6%
WinoGrande—77.2%

Math Not comparable

Grok Build 0.1: —, Mixtral 8x7B: 18.8 (#289)

Math benchmarks
BenchmarkGrok Build 0.1Mixtral 8x7B
Omni-MATH—10.5%
LMArena Math—1147
MATH Level 5—10%
GSM8K—74.4%

Knowledge Not comparable

Grok Build 0.1: —, Mixtral 8x7B: 11.0 (#301)

Knowledge benchmarks
BenchmarkGrok Build 0.1Mixtral 8x7B
GPQA Diamond—30.6%
MMLU-Pro—33.5%
GPQA (HELM)—29.6%
LMArena Expert—1088
ARC (AI2) Challenge—87.3%
MMLU—70.6%
OpenBookQA—85.8%
TriviaQA—82.2%

Multilingual Not comparable

Grok Build 0.1: —, Mixtral 8x7B: 29.6 (#266)

Multilingual benchmarks
BenchmarkGrok Build 0.1Mixtral 8x7B
LMArena Non-English—1077
LMArena Chinese—1055
LMArena French—1166
LMArena German—1114
LMArena Japanese—931
LMArena Korean—968
LMArena Russian—1090
LMArena Spanish—1111

Instruction Following Not comparable

Grok Build 0.1: —, Mixtral 8x7B: 51.0 (#297)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Mixtral 8x7B
IFEval—57.5%
LMArena Instruction Following—1109

Long Context Not comparable

Grok Build 0.1: —, Mixtral 8x7B: 33.4 (#260)

Long Context benchmarks
BenchmarkGrok Build 0.1Mixtral 8x7B
LMArena Longer Query—1103

Writing & Preference Not comparable

Grok Build 0.1: —, Mixtral 8x7B: 34.2 (#270)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Mixtral 8x7B
LMArena Text—1132
LMArena Creative Writing—1109
WildBench—67.3%
LMArena Multi-Turn—1115

Frequently asked questions

Is Grok Build 0.1 better than Mixtral 8x7B?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 27.1 on the Noometry Index. Mixtral 8x7B costs 1.8× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Which is cheaper, Grok Build 0.1 or Mixtral 8x7B?

Mixtral 8x7B is cheaper. It lists at $0.70 per million input tokens and $0.70 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or Mixtral 8x7B better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 32.8 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 32K.

How many benchmarks do Grok Build 0.1 and Mixtral 8x7B share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Mixtral 8x7B has 38.

Related comparisons

Go deeper