Model comparison

Chatgpt 4o Latest 20250326 vs Mistral Large 4

Chatgpt 4o Latest 20250326 and Mistral Large 4 score almost the same on the Noometry Index (43.8 vs 43.1), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Mistral Large 4 Mistral AI

43.1

Rank #99 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 4 categories and Mistral Large 4 in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Chatgpt 4o Latest 20250326 leads 33.8 to 22.5.

Side by side

Chatgpt 4o Latest 20250326 and Mistral Large 4 specifications
Chatgpt 4o Latest 20250326Mistral Large 4
ProviderOpenAIMistral AI
Noometry Index43.843.1
Released—2026-10-06
WeightsProprietaryProprietary
Context window—1.05M
Max output—262K
Input $ / M tokens—$0.68
Output $ / M tokens—$2.09
Results tracked2115

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large 4 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Mistral Large 4: 48.6 (#57)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 4
LMArena Coding14131475
LMArena WebDev—1541

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Mistral Large 4: 22.5 (#192)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 4
LMArena Hard Prompts14241444
Kagi LLM Benchmark75%—
NYT Connections (extended)—27.4%

Math Mistral Large 4 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Mistral Large 4: 40.4 (#91)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 4
LMArena Math14071488

Knowledge Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Mistral Large 4: 36.6 (#166)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 4
LMArena Expert14011447
SimpleQA Verified—20%
Confabulations16.6%—

Multimodal Not comparable

Chatgpt 4o Latest 20250326: 39.6 (#58), Mistral Large 4: —

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 4
LMArena Vision1243—

Multilingual Too close to call

Chatgpt 4o Latest 20250326: 52.9 (#76), Mistral Large 4: 52.6 (#82)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 4
LMArena Non-English14191415
LMArena Chinese14571491
LMArena Russian14291414
LMArena French1446—
LMArena German1423—
LMArena Japanese1405—
LMArena Korean1396—
LMArena Spanish1434—

Instruction Following Mistral Large 4 leads

Chatgpt 4o Latest 20250326: 74.0 (#107), Mistral Large 4: 75.0 (#76)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 4
LMArena Instruction Following14031424

Long Context Too close to call

Chatgpt 4o Latest 20250326: 43.1 (#107), Mistral Large 4: 43.6 (#89)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 4
LMArena Longer Query14131429

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Mistral Large 4: 60.4 (#97)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Large 4
LMArena Text14291427
LMArena Creative Writing14051361
LMArena Multi-Turn14541424
EQ-Bench Creative Writing1501—

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Mistral Large 4?

Chatgpt 4o Latest 20250326 and Mistral Large 4 score almost the same on the Noometry Index (43.8 vs 43.1), so choose on price, context window or the category you care about most.

Is Chatgpt 4o Latest 20250326 or Mistral Large 4 better for coding?

Mistral Large 4 scores higher on coding benchmarks: 48.6 versus 41.6 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Mistral Large 4 share?

12 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Mistral Large 4 has 15.

Related comparisons

Go deeper