Model comparison

Grok 4.3 vs Mistral Large

Grok 4.3 is the stronger model overall, scoring 43.8 to 31.9 on the Noometry Index.

Last verified . 26 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Grok 4.3 scores higher in 8 categories and Mistral Large in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.3 leads 46.0 to 18.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 93.3% for Grok 4.3 and 8.5% for Mistral Large.
  • Grok 4.3 is cheaper at $1.25 / $2.50 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Grok 4.3 accepts more context: 1M tokens versus 131K.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Grok 4.3 and Mistral Large specifications
Grok 4.3Mistral Large
ProviderxAIMistral AI
Noometry Index43.831.9
Released2026-04-172024-02-26
WeightsProprietaryOpen
Context window1M131K
Max output30K16K
Input $ / M tokens$1.25$2
Output $ / M tokens$2.50$6
Results tracked4051

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

Grok 4.3: 41.6 (#121), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGrok 4.3Mistral Large
SciCode47.3%36.2%
LMArena Coding14151277
ALE-Bench944.17264.7
LMArena WebDev1357—
WeirdML49.9%—
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Too close to call

Grok 4.3: 27.7 (#99), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Mistral Large
Berkeley Function Calling Leaderboard—38.4%
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGrok 4.3Mistral Large
CritPt8%0%
LMArena Hard Prompts13961257
DTBench90.7%65.1%
LMCA38.3%16.7%
Epoch Capabilities Index149.16128.52
ForecastBench60.357.1
SimpleBench—22.5%
NYT Connections (extended)55.2%—
Chess Puzzles25%—
LiveBench Reasoning—43.5%
LiveBench Data Analysis—50.1%
LiveBench—48.4%

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGrok 4.3Mistral Large
OTIS Mock AIME 2024-202593.3%8.5%
LMArena Math13881262
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
ProofBench11%—
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGrok 4.3Mistral Large
GPQA Diamond88.8%51.3%
LMArena Expert13851232
SimpleQA Verified33.2%—
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
MMLU—80%

Multimodal Not comparable

Grok 4.3: 31.6 (#104), Mistral Large: —

Multimodal benchmarks
BenchmarkGrok 4.3Mistral Large
LMArena Vision1229—
Blueprint-Bench 20%—

Multilingual Grok 4.3 leads

Grok 4.3: 50.5 (#120), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGrok 4.3Mistral Large
LMArena Non-English13851237
LMArena Chinese14221240
LMArena French14121325
LMArena German13951254
LMArena Japanese13791188
LMArena Korean13561202
LMArena Russian13991257
LMArena Spanish13981268

Instruction Following Grok 4.3 leads

Grok 4.3: 72.1 (#140), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGrok 4.3Mistral Large
LMArena Instruction Following13661249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Grok 4.3 leads

Grok 4.3: 42.5 (#123), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGrok 4.3Mistral Large
LMArena Longer Query13931261

Writing & Preference Grok 4.3 leads

Grok 4.3: 58.5 (#118), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGrok 4.3Mistral Large
LMArena Text13971266
LMArena Creative Writing13801243
LMArena Multi-Turn14061260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
EQ-Bench 41075—
LiveBench Language—39.4%

Frequently asked questions

Is Grok 4.3 better than Mistral Large?

Grok 4.3 is the stronger model overall, scoring 43.8 to 31.9 on the Noometry Index.

Which is cheaper, Grok 4.3 or Mistral Large?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; Mistral Large lists at $2 and $6.

Is Grok 4.3 or Mistral Large better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 131K.

How many benchmarks do Grok 4.3 and Mistral Large share?

26 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper