Model comparison

Chatgpt 4o Latest 20250326 vs Mistral 7B

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 23.0 on the Noometry Index.

Last verified . 16 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 8 categories and Mistral 7B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Chatgpt 4o Latest 20250326 leads 62.6 to 30.7.
  • Mistral 7B has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and Mistral 7B specifications
Chatgpt 4o Latest 20250326Mistral 7B
ProviderOpenAIMistral AI
Noometry Index43.823.0
Released—2023-09-27
WeightsProprietaryOpen
Context window—8K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.25
Results tracked2137

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Mistral 7B: 26.4 (#326)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral 7B
LMArena Coding14131082
BigCodeBench Instruct—19.5%
BigCodeBench Complete—27.3%
HumanEval+—36%
MBPP+—42.1%

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Mistral 7B: 13.1 (#336)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral 7B
LMArena Hard Prompts14241067
Kagi LLM Benchmark75%—
Chess Puzzles—0%
DTBench—42.5%
Adversarial NLI—47.1%
BIG-Bench Hard—56.1%
Epoch Capabilities Index—112.21
HellaSwag—81%
PIQA—83%
WinoGrande—75.3%

Math Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Mistral 7B: 8.1 (#325)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral 7B
LMArena Math14071085
OTIS Mock AIME 2024-2025—0.3%
MATH Level 5—3.7%
GSM8K—54.4%

Knowledge Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Mistral 7B: 7.4 (#311)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral 7B
LMArena Expert14011036
GPQA Diamond—15.2%
Confabulations16.6%—
ARC (AI2) Challenge—78.6%
BoolQ—87.4%
MMLU—62.5%
OpenBookQA—79.8%
TriviaQA—75.2%

Multimodal Not comparable

Chatgpt 4o Latest 20250326: 39.6 (#58), Mistral 7B: —

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral 7B
LMArena Vision1243—

Multilingual Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 52.9 (#76), Mistral 7B: 25.8 (#283)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral 7B
LMArena Non-English14191012
LMArena Chinese14571009
LMArena French14461037
LMArena German1423987
LMArena Japanese1405878
LMArena Russian14291018
LMArena Spanish14341026
LMArena Korean1396—

Instruction Following Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 74.0 (#107), Mistral 7B: 54.2 (#280)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral 7B
LMArena Instruction Following14031060

Long Context Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 43.1 (#107), Mistral 7B: 32.2 (#271)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral 7B
LMArena Longer Query14131060

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Mistral 7B: 30.7 (#286)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Mistral 7B
LMArena Text14291090
LMArena Creative Writing14051068
LMArena Multi-Turn14541062
EQ-Bench Creative Writing1501—

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Mistral 7B?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 23.0 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Mistral 7B better for coding?

Chatgpt 4o Latest 20250326 scores higher on coding benchmarks: 41.6 versus 26.4 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Mistral 7B share?

16 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Mistral 7B has 37.

Related comparisons

Go deeper