Model comparison

Chatgpt 4o Latest 20250326 vs Mistral Medium

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 36.3 on the Noometry Index.

Last verified . 19 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 9 categories and Mistral Medium in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Chatgpt 4o Latest 20250326 leads 39.2 to 25.0.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 75% for Chatgpt 4o Latest 20250326 and 50% for Mistral Medium.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and Mistral Medium specifications
Chatgpt 4o Latest 20250326Mistral Medium
ProviderOpenAIMistral AI
Noometry Index43.836.3
Released—2023-12-11
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked2136

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Medium
LMArena Coding14131434
FrontierCode—8%
SciCode—40.2%
WeirdML—43.7%
ALE-Bench—763.98

Agentic & Tool Use Not comparable

Chatgpt 4o Latest 20250326: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Medium
Kagi LLM Benchmark75%50%
LMArena Hard Prompts14241426
CritPt—0%
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Medium
LMArena Math14071408
OTIS Mock AIME 2024-2025—32.2%
ProofBench—9%
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Medium
LMArena Expert14011408
GPQA Diamond—59.5%
Humanity's Last Exam—4.5%
Confabulations16.6%—
Vectara Hallucination Rate—22.7%

Multimodal Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.6 (#58), Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Medium
LMArena Vision12431172

Multilingual Too close to call

Chatgpt 4o Latest 20250326: 52.9 (#76), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Medium
LMArena Non-English14191408
LMArena Chinese14571447
LMArena French14461459
LMArena German14231432
LMArena Japanese14051378
LMArena Korean13961380
LMArena Russian14291411
LMArena Spanish14341433

Instruction Following Too close to call

Chatgpt 4o Latest 20250326: 74.0 (#107), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Medium
LMArena Instruction Following14031398

Long Context Too close to call

Chatgpt 4o Latest 20250326: 43.1 (#107), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Medium
LMArena Longer Query14131406

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Medium
LMArena Text14291424
LMArena Creative Writing14051391
LMArena Multi-Turn14541418
Short-Story Creative Writing—77.3%
EQ-Bench Creative Writing1501—

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Mistral Medium?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 36.3 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Mistral Medium better for coding?

Chatgpt 4o Latest 20250326 scores higher on coding benchmarks: 41.6 versus 34.2 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Mistral Medium share?

19 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper