Model comparison

GPT-4 vs Mistral

GPT-4 and Mistral score almost the same on the Noometry Index (29.1 vs 29.9), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Mistral Mistral AI

29.9

Rank #303 Confirmed

Summary

  • They share 17 benchmarks with published results for both. GPT-4 scores higher in 4 categories and Mistral in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where GPT-4 leads 65.3 to 52.6.

Side by side

GPT-4 and Mistral specifications
GPT-4Mistral
ProviderOpenAIMistral AI
Noometry Index29.129.9
Released2023-03-14—
WeightsProprietaryProprietary
Context window8K—
Max output8K—
Input $ / M tokens$30—
Output $ / M tokens$60—
Results tracked3822

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

GPT-4: 31.6 (#283), Mistral: 33.8 (#250)

Coding benchmarks
BenchmarkGPT-4Mistral
LMArena Coding12541162
WeirdML12.4%—
BigCodeBench Instruct46%—
BigCodeBench Complete57.2%—
HumanEval+79.3%—

Agentic & Tool Use Not comparable

GPT-4: —, Mistral: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4Mistral
METR Time Horizons36.1%—

Reasoning Mistral leads

GPT-4: 17.8 (#289), Mistral: 22.2 (#200)

Reasoning benchmarks
BenchmarkGPT-4Mistral
LMArena Hard Prompts12411149
Chess Puzzles4%—
Mystery Game Puzzles12%—
DTBench62.7%—
LMCA17.1%—
BIG-Bench Hard75.1%—
Epoch Capabilities Index125.89—
ForecastBench57.8—
HellaSwag95.3%—
WinoGrande87.5%—

Math Mistral leads

GPT-4: 10.8 (#309), Mistral: 22.3 (#278)

Math benchmarks
BenchmarkGPT-4Mistral
LMArena Math12691180
OTIS Mock AIME 2024-20251.1%—
Omni-MATH—7.2%
MATH Level 523%—
GSM8K92%—

Knowledge GPT-4 leads

GPT-4: 18.4 (#282), Mistral: 16.6 (#288)

Knowledge benchmarks
BenchmarkGPT-4Mistral
LMArena Expert12111125
GPQA Diamond35.7%—
MMLU-Pro—27.7%
GPQA (HELM)—30.3%
MMLU86.4%—
TriviaQA84.8%—

Multilingual GPT-4 leads

GPT-4: 40.6 (#215), Mistral: 32.8 (#254)

Multilingual benchmarks
BenchmarkGPT-4Mistral
LMArena Non-English12461129
LMArena Chinese12421109
LMArena French12831180
LMArena German12511155
LMArena Japanese12091013
LMArena Korean11841032
LMArena Russian12511168
LMArena Spanish12611143

Instruction Following GPT-4 leads

GPT-4: 65.3 (#222), Mistral: 52.6 (#288)

Instruction Following benchmarks
BenchmarkGPT-4Mistral
LMArena Instruction Following12411152
IFEval—56.8%

Long Context GPT-4 leads

GPT-4: 37.7 (#212), Mistral: 35.0 (#245)

Long Context benchmarks
BenchmarkGPT-4Mistral
LMArena Longer Query12441153

Writing & Preference Mistral leads

GPT-4: 34.9 (#268), Mistral: 37.0 (#260)

Writing & Preference benchmarks
BenchmarkGPT-4Mistral
LMArena Text12631165
LMArena Creative Writing12441158
LMArena Multi-Turn12571147
EQ-Bench Creative Writing752—
WildBench—66%

Frequently asked questions

Is GPT-4 better than Mistral?

GPT-4 and Mistral score almost the same on the Noometry Index (29.1 vs 29.9), so choose on price, context window or the category you care about most.

Is GPT-4 or Mistral better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 31.6 in the Noometry coding category.

How many benchmarks do GPT-4 and Mistral share?

17 benchmarks have published results for both models. GPT-4 has 38 scored results on Noometry and Mistral has 22.

Related comparisons

Go deeper