Model comparison

Chatgpt 4o Latest 20250326 vs Mistral Large 3

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 39.1 on the Noometry Index.

Last verified . 20 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 6 categories and Mistral Large 3 in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Chatgpt 4o Latest 20250326 leads 33.8 to 15.2.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 75% for Chatgpt 4o Latest 20250326 and 50.9% for Mistral Large 3.
  • Mistral Large 3 has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and Mistral Large 3 specifications
Chatgpt 4o Latest 20250326Mistral Large 3
ProviderOpenAIMistral AI
Noometry Index43.839.1
Released—2025-12-02
WeightsProprietaryOpen
Context window—262K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.75
Results tracked2124

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Mistral Large 3: 34.4 (#237)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 3
LMArena Coding14131448
LMArena WebDev—1230

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Mistral Large 3: 15.2 (#319)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 3
Kagi LLM Benchmark75%50.9%
LMArena Hard Prompts14241429
NYT Connections (extended)—7.5%
Thematic Generalization—23%

Math Too close to call

Chatgpt 4o Latest 20250326: 38.6 (#134), Mistral Large 3: 38.7 (#129)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 3
LMArena Math14071414

Knowledge Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Mistral Large 3: 36.0 (#177)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 3
LMArena Expert14011421
Confabulations16.6%—
Vectara Hallucination Rate—14.5%

Multimodal Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.6 (#58), Mistral Large 3: 38.2 (#66)

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 3
LMArena Vision12431221

Multilingual Too close to call

Chatgpt 4o Latest 20250326: 52.9 (#76), Mistral Large 3: 52.5 (#84)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 3
LMArena Non-English14191413
LMArena Chinese14571447
LMArena French14461455
LMArena German14231437
LMArena Japanese14051394
LMArena Korean13961384
LMArena Russian14291411
LMArena Spanish14341440

Instruction Following Too close to call

Chatgpt 4o Latest 20250326: 74.0 (#107), Mistral Large 3: 74.0 (#108)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 3
LMArena Instruction Following14031403

Long Context Too close to call

Chatgpt 4o Latest 20250326: 43.1 (#107), Mistral Large 3: 43.1 (#105)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 3
LMArena Longer Query14131413

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Mistral Large 3: 60.0 (#101)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 3
LMArena Text14291428
LMArena Creative Writing14051386
EQ-Bench Creative Writing15011412
LMArena Multi-Turn14541429

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Mistral Large 3?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 39.1 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Mistral Large 3 better for coding?

Chatgpt 4o Latest 20250326 scores higher on coding benchmarks: 41.6 versus 34.4 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Mistral Large 3 share?

20 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Mistral Large 3 has 24.

Related comparisons

Go deeper