Model comparison

Chatgpt 4o Latest 20250326 vs Mistral Small

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 33.4 on the Noometry Index.

Last verified . 19 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 9 categories and Mistral Small in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Chatgpt 4o Latest 20250326 leads 38.6 to 16.4.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 75% for Chatgpt 4o Latest 20250326 and 37.8% for Mistral Small.
  • Mistral Small has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and Mistral Small specifications
Chatgpt 4o Latest 20250326Mistral Small
ProviderOpenAIMistral AI
Noometry Index43.833.4
Released—2024-02-26
WeightsProprietaryOpen
Context window—262K
Max output—256K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.60
Results tracked2139

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small
LMArena Coding14131362
SciCode—26.5%
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
BigCodeBench Complete—46.6%
ALE-Bench—497.62

Agentic & Tool Use Not comparable

Chatgpt 4o Latest 20250326: —, Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small
Berkeley Function Calling Leaderboard—37.1%

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small
Kagi LLM Benchmark75%37.8%
LMArena Hard Prompts14241335
CritPt—0%
LiveBench Reasoning—44.8%
DTBench—70.9%
LiveBench Data Analysis—53.7%
LMCA—20.6%
LiveBench—44%

Math Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small
LMArena Math14071341
OTIS Mock AIME 2024-2025—5.8%
LiveBench Math—39.9%
MATH Level 5—46.8%

Knowledge Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small
LMArena Expert14011291
GPQA Diamond—47.5%
Confabulations16.6%—
Vectara Hallucination Rate—5.1%
MMLU—68.7%

Multimodal Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.6 (#58), Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small
LMArena Vision12431142

Multilingual Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 52.9 (#76), Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small
LMArena Non-English14191315
LMArena Chinese14571340
LMArena French14461337
LMArena German14231340
LMArena Japanese14051275
LMArena Korean13961259
LMArena Russian14291324
LMArena Spanish14341346

Instruction Following Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 74.0 (#107), Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small
LMArena Instruction Following14031310
LiveBench Instruction Following—63.7%

Long Context Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 43.1 (#107), Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small
LMArena Longer Query14131327

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral Small
LMArena Text14291338
LMArena Creative Writing14051305
LMArena Multi-Turn14541344
EQ-Bench Creative Writing1501—
LiveBench Language—30.5%

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Mistral Small?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 33.4 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Mistral Small better for coding?

Chatgpt 4o Latest 20250326 scores higher on coding benchmarks: 41.6 versus 34.0 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Mistral Small share?

19 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper