Model comparison

Chatgpt 4o Latest 20250326 vs Mistral Small 3.1

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 31.7 on the Noometry Index.

Last verified . 19 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 9 categories and Mistral Small 3.1 in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Chatgpt 4o Latest 20250326 leads 62.6 to 37.0.
  • Mistral Small 3.1 has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and Mistral Small 3.1 specifications
Chatgpt 4o Latest 20250326Mistral Small 3.1
ProviderOpenAIMistral AI
Noometry Index43.831.7
Released—2025-03-17
WeightsProprietaryOpen
Context window—128K
Max output—102K
Input $ / M tokens—$0.35
Output $ / M tokens—$0.56
Results tracked2128

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3.1
LMArena Coding14131309

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3.1
LMArena Hard Prompts14241278
Kagi LLM Benchmark75%—
Chess Puzzles—1%
Epoch Capabilities Index—127.48

Math Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3.1
LMArena Math14071262
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%

Knowledge Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3.1
LMArena Expert14011257
GPQA Diamond—41.9%
MMLU-Pro—61%
Confabulations16.6%—
GPQA (HELM)—39.2%

Multimodal Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.6 (#58), Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3.1
LMArena Vision12431136

Multilingual Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 52.9 (#76), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3.1
LMArena Non-English14191255
LMArena Chinese14571253
LMArena French14461273
LMArena German14231266
LMArena Japanese14051208
LMArena Korean13961206
LMArena Russian14291263
LMArena Spanish14341283

Instruction Following Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 74.0 (#107), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3.1
LMArena Instruction Following14031264
IFEval—75%

Long Context Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 43.1 (#107), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3.1
LMArena Longer Query14131299

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3.1
LMArena Text14291277
LMArena Creative Writing14051253
EQ-Bench Creative Writing1501761
LMArena Multi-Turn14541270
WildBench—78.8%

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Mistral Small 3.1?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 31.7 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Mistral Small 3.1 better for coding?

Chatgpt 4o Latest 20250326 scores higher on coding benchmarks: 41.6 versus 38.3 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Mistral Small 3.1 share?

19 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper