Model comparison

Grok 4.20 (Non-Reasoning) vs Mixtral 8x22B

Grok 4.20 (Non-Reasoning) is the stronger model overall, scoring 48.6 to 27.1 on the Noometry Index.

Last verified . 22 shared benchmarks.

Grok 4.20 (Non-Reasoning) xAI

48.6

Rank #54 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Grok 4.20 (Non-Reasoning) scores higher in 9 categories and Mixtral 8x22B in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.20 (Non-Reasoning) leads 52.8 to 15.1.
  • The biggest single-benchmark swing is GPQA Diamond: 89.3% for Grok 4.20 (Non-Reasoning) and 34.1% for Mixtral 8x22B.
  • Grok 4.20 (Non-Reasoning) is cheaper at $1.25 / $2.50 per million input/output tokens, against $2 / $6 for Mixtral 8x22B.
  • Grok 4.20 (Non-Reasoning) accepts more context: 1M tokens versus 64K.
  • Mixtral 8x22B has downloadable open weights; the other is API-only.

Side by side

Grok 4.20 (Non-Reasoning) and Mixtral 8x22B specifications
Grok 4.20 (Non-Reasoning)Mixtral 8x22B
ProviderxAIMistral AI
Noometry Index48.627.1
Released2026-02-172024-04-17
WeightsProprietaryOpen
Context window1M64K
Max output30K64K
Input $ / M tokens$1.25$2
Output $ / M tokens$2.50$6
Results tracked4634

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 42.1 (#112), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Mixtral 8x22B
WeirdML52.3%3.2%
LMArena Coding14591166
LMArena WebDev1375—
BigCodeBench Instruct—40.6%
BigCodeBench Complete—50.2%
ALE-Bench1,150—
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 34.4 (#46), Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Mixtral 8x22B
Terminal-Bench57.3%—
τ²-bench Banking18%—
Cybench—7.5%
LMArena Search1189—
Vending-Bench 24,663—

Reasoning Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 52.3 (#32), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Mixtral 8x22B
LMArena Hard Prompts14511150
DTBench90.1%55.1%
Epoch Capabilities Index151.98122.03
ForecastBench61.456.3
ARC-AGI-265.1%—
Kagi LLM Benchmark75%—
NYT Connections (extended)85.4%—
ARC-AGI-189.5%—
Chess Puzzles24%—
Thematic Generalization63.8%—
LMCA38.7%—

Math Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 48.2 (#65), Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Mixtral 8x22B
LMArena Math14551184
FrontierMath (Tiers 1-3)44.9%—
FrontierMath Tier 417.1%—
OTIS Mock AIME 2024-202592.2%—
ProofBench14%—
Omni-MATH—16.3%
MATH Level 5—24.2%

Knowledge Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 52.8 (#60), Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Mixtral 8x22B
GPQA Diamond89.3%34.1%
LMArena Expert14391113
SimpleQA Verified30.2%—
MMLU-Pro—46%
GPQA (HELM)—33.4%
MMLU—77.8%

Multimodal Not comparable

Grok 4.20 (Non-Reasoning): 33.3 (#98), Mixtral 8x22B: —

Multimodal benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Mixtral 8x22B
LMArena Vision1263—
Blueprint-Bench 20%—
LMArena Document1416—

Multilingual Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 54.5 (#40), Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Mixtral 8x22B
LMArena Non-English14411128
LMArena Chinese14811116
LMArena French14761166
LMArena German14651141
LMArena Japanese14491037
LMArena Korean14171057
LMArena Russian14581158
LMArena Spanish14431151

Instruction Following Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 74.8 (#83), Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Mixtral 8x22B
LMArena Instruction Following14201147
IFEval—72.4%

Long Context Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 45.5 (#34), Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Mixtral 8x22B
LMArena Longer Query14371144
CL-bench22.2%—
CL-bench Life11.9%—

Writing & Preference Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 65.7 (#44), Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Mixtral 8x22B
LMArena Text14511162
LMArena Creative Writing14381141
LMArena Multi-Turn14561130
EQ-Bench Creative Writing1574—
WildBench—71.1%

Frequently asked questions

Is Grok 4.20 (Non-Reasoning) better than Mixtral 8x22B?

Grok 4.20 (Non-Reasoning) is the stronger model overall, scoring 48.6 to 27.1 on the Noometry Index.

Which is cheaper, Grok 4.20 (Non-Reasoning) or Mixtral 8x22B?

Grok 4.20 (Non-Reasoning) is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; Mixtral 8x22B lists at $2 and $6.

Is Grok 4.20 (Non-Reasoning) or Mixtral 8x22B better for coding?

Grok 4.20 (Non-Reasoning) scores higher on coding benchmarks: 42.1 versus 24.2 in the Noometry coding category.

Which has the bigger context window?

Grok 4.20 (Non-Reasoning) does, with 1M tokens against 64K.

How many benchmarks do Grok 4.20 (Non-Reasoning) and Mixtral 8x22B share?

22 benchmarks have published results for both models. Grok 4.20 (Non-Reasoning) has 46 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper