Model comparison

GPT-4 vs Mistral Medium

Mistral Medium is the stronger model overall, scoring 36.3 to 29.1 on the Noometry Index.

Last verified . 23 shared benchmarks.

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 23 benchmarks with published results for both. GPT-4 scores higher in 0 categories and Mistral Medium in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 34.9.
  • The biggest single-benchmark swing is MATH Level 5: 23% for GPT-4 and 81.6% for Mistral Medium.
  • Mistral Medium is cheaper at $1.50 / $7.50 per million input/output tokens, against $30 / $60 for GPT-4.
  • Mistral Medium accepts more context: 262K tokens versus 8K.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

GPT-4 and Mistral Medium specifications
GPT-4Mistral Medium
ProviderOpenAIMistral AI
Noometry Index29.136.3
Released2023-03-142023-12-11
WeightsProprietaryOpen
Context window8K262K
Max output8K262K
Input $ / M tokens$30$1.50
Output $ / M tokens$60$7.50
Results tracked3836

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium leads

GPT-4: 31.6 (#283), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkGPT-4Mistral Medium
WeirdML12.4%43.7%
LMArena Coding12541434
FrontierCode—8%
SciCode—40.2%
BigCodeBench Instruct46%—
BigCodeBench Complete57.2%—
ALE-Bench—763.98
HumanEval+79.3%—

Agentic & Tool Use Not comparable

GPT-4: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkGPT-4Mistral Medium
Berkeley Function Calling Leaderboard—37.7%
METR Time Horizons36.1%—

Reasoning Mistral Medium leads

GPT-4: 17.8 (#289), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkGPT-4Mistral Medium
LMArena Hard Prompts12411426
DTBench62.7%75.5%
LMCA17.1%26.1%
Kagi LLM Benchmark—50%
CritPt—0%
Chess Puzzles4%—
Mystery Game Puzzles12%—
Surface Evolver Bench—26.9%
BIG-Bench Hard75.1%—
Epoch Capabilities Index125.89—
ForecastBench57.8—
HellaSwag95.3%—
WinoGrande87.5%—

Math Mistral Medium leads

GPT-4: 10.8 (#309), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkGPT-4Mistral Medium
OTIS Mock AIME 2024-20251.1%32.2%
LMArena Math12691408
MATH Level 523%81.6%
ProofBench—9%
FrontierMath (Feb 2025 set)—0.3%
GSM8K92%—

Knowledge Mistral Medium leads

GPT-4: 18.4 (#282), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkGPT-4Mistral Medium
GPQA Diamond35.7%59.5%
LMArena Expert12111408
Humanity's Last Exam—4.5%
Vectara Hallucination Rate—22.7%
MMLU86.4%—
TriviaQA84.8%—

Multimodal Not comparable

GPT-4: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkGPT-4Mistral Medium
LMArena Vision—1172

Multilingual Mistral Medium leads

GPT-4: 40.6 (#215), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkGPT-4Mistral Medium
LMArena Non-English12461408
LMArena Chinese12421447
LMArena French12831459
LMArena German12511432
LMArena Japanese12091378
LMArena Korean11841380
LMArena Russian12511411
LMArena Spanish12611433

Instruction Following Mistral Medium leads

GPT-4: 65.3 (#222), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkGPT-4Mistral Medium
LMArena Instruction Following12411398

Long Context Mistral Medium leads

GPT-4: 37.7 (#212), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkGPT-4Mistral Medium
LMArena Longer Query12441406

Writing & Preference Mistral Medium leads

GPT-4: 34.9 (#268), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkGPT-4Mistral Medium
LMArena Text12631424
LMArena Creative Writing12441391
LMArena Multi-Turn12571418
Short-Story Creative Writing—77.3%
EQ-Bench Creative Writing752—

Frequently asked questions

Is GPT-4 better than Mistral Medium?

Mistral Medium is the stronger model overall, scoring 36.3 to 29.1 on the Noometry Index.

Which is cheaper, GPT-4 or Mistral Medium?

Mistral Medium is cheaper. It lists at $1.50 per million input tokens and $7.50 per million output tokens; GPT-4 lists at $30 and $60.

Is GPT-4 or Mistral Medium better for coding?

Mistral Medium scores higher on coding benchmarks: 34.2 versus 31.6 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 8K.

How many benchmarks do GPT-4 and Mistral Medium share?

23 benchmarks have published results for both models. GPT-4 has 38 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper