Model comparison

Chatgpt 4o Latest 20250326 vs Mistral Small 3

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 31.2 on the Noometry Index.

Last verified . 18 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Mistral Small 3 Mistral AI

31.2

Rank #278 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 8 categories and Mistral Small 3 in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Chatgpt 4o Latest 20250326 leads 62.6 to 32.2.
  • The biggest single-benchmark swing is Confabulations: 16.6% for Chatgpt 4o Latest 20250326 and 25.2% for Mistral Small 3.
  • Mistral Small 3 has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and Mistral Small 3 specifications
Chatgpt 4o Latest 20250326Mistral Small 3
ProviderOpenAIMistral AI
Noometry Index43.831.2
Released—2025-01-30
WeightsProprietaryOpen
Context window—33K
Max output—16K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.08
Results tracked2124

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Mistral Small 3: 36.5 (#207)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3
LMArena Coding14131246
BigCodeBench Instruct—45.3%
BigCodeBench Complete—50.4%

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Mistral Small 3: 18.9 (#273)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3
LMArena Hard Prompts14241233
Kagi LLM Benchmark75%—
Chess Puzzles—0%
Epoch Capabilities Index—127.07

Math Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Mistral Small 3: 16.3 (#295)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3
LMArena Math14071240
OTIS Mock AIME 2024-2025—6.7%

Knowledge Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Mistral Small 3: 25.1 (#263)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3
Confabulations16.6%25.2%
LMArena Expert14011202
GPQA Diamond—47.3%

Multimodal Not comparable

Chatgpt 4o Latest 20250326: 39.6 (#58), Mistral Small 3: —

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3
LMArena Vision1243—

Multilingual Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 52.9 (#76), Mistral Small 3: 37.3 (#236)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3
LMArena Non-English14191198
LMArena Chinese14571204
LMArena French14461203
LMArena German14231211
LMArena Japanese14051111
LMArena Korean13961188
LMArena Russian14291216
LMArena Spanish1434—

Instruction Following Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 74.0 (#107), Mistral Small 3: 63.7 (#229)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3
LMArena Instruction Following14031214

Long Context Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 43.1 (#107), Mistral Small 3: 37.8 (#211)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3
LMArena Longer Query14131246

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Mistral Small 3: 32.2 (#280)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small 3
LMArena Text14291234
LMArena Creative Writing14051195
EQ-Bench Creative Writing1501707
LMArena Multi-Turn14541217

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Mistral Small 3?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 31.2 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Mistral Small 3 better for coding?

Chatgpt 4o Latest 20250326 scores higher on coding benchmarks: 41.6 versus 36.5 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Mistral Small 3 share?

18 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Mistral Small 3 has 24.

Related comparisons

Go deeper