Model comparison

GPT-4o mini vs Mistral Small

Mistral Small is the stronger model overall, scoring 33.4 to 25.5 on the Noometry Index.

Last verified . 34 shared benchmarks.

GPT-4o mini OpenAI

25.5

Rank #343 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 34 benchmarks with published results for both. GPT-4o mini scores higher in 0 categories and Mistral Small in 10 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Small leads 31.0 to 17.7.
  • The biggest single-benchmark swing is DTBench: 54.4% for GPT-4o mini and 70.9% for Mistral Small.
  • Both cost about the same: $0.15 input and $0.60 output per million tokens.
  • Mistral Small accepts more context: 262K tokens versus 128K.
  • Mistral Small has downloadable open weights; the other is API-only.

Side by side

GPT-4o mini and Mistral Small specifications
GPT-4o miniMistral Small
ProviderOpenAIMistral AI
Noometry Index25.533.4
Released2024-07-182024-02-26
WeightsProprietaryOpen
Context window128K262K
Max output16K256K
Input $ / M tokens$0.15$0.15
Output $ / M tokens$0.60$0.60
Results tracked6039

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small leads

GPT-4o mini: 22.0 (#335), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkGPT-4o miniMistral Small
BigCodeBench Instruct46.1%36.1%
LiveBench Coding43.1%36.2%
LMArena Coding12901362
BigCodeBench Complete57.4%46.6%
Aider Polyglot3.6%—
SciCode—26.5%
WeirdML11.8%—
ALE-Bench—497.62
HumanEval+83.5%—
MBPP+72.2%—

Agentic & Tool Use Too close to call

GPT-4o mini: 27.5 (#101), Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkGPT-4o miniMistral Small
Berkeley Function Calling Leaderboard—37.1%
BALROG17.4%—

Reasoning Mistral Small leads

GPT-4o mini: 8.7 (#347), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkGPT-4o miniMistral Small
Kagi LLM Benchmark28.8%37.8%
LiveBench Reasoning32.8%44.8%
LMArena Hard Prompts12671335
DTBench54.4%70.9%
LiveBench Data Analysis50%53.7%
LMCA10.4%20.6%
LiveBench41.3%44%
ARC-AGI-20%—
SimpleBench10.7%—
CritPt—0%
Chess Puzzles0%—
Mystery Game Puzzles12%—
Epoch Capabilities Index126.56—
PIQA88.7%—

Math Mistral Small leads

GPT-4o mini: 10.4 (#314), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkGPT-4o miniMistral Small
OTIS Mock AIME 2024-20256.9%5.8%
LiveBench Math36.3%39.9%
LMArena Math12671341
MATH Level 552.6%46.8%
FrontierMath (Tiers 1-3)0.7%—
Omni-MATH28%—
GSM8K91.3%—

Knowledge Mistral Small leads

GPT-4o mini: 17.7 (#284), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkGPT-4o miniMistral Small
GPQA Diamond37.7%47.5%
LMArena Expert12351291
MMLU81.8%68.7%
SimpleQA Verified8.3%—
MMLU-Pro60.3%—
Confabulations37.2%—
Vectara Hallucination Rate—5.1%
GPQA (HELM)36.8%—
BoolQ88.7%—

Multimodal Mistral Small leads

GPT-4o mini: 25.9 (#122), Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkGPT-4o miniMistral Small
LMArena Vision10661142
Video-MME64.8%—
GeoBench64%—
VPCT34%—

Multilingual Mistral Small leads

GPT-4o mini: 42.0 (#199), Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkGPT-4o miniMistral Small
LMArena Non-English12661315
LMArena Chinese12651340
LMArena French12971337
LMArena German12721340
LMArena Japanese12161275
LMArena Korean11951259
LMArena Russian12751324
LMArena Spanish12761346

Instruction Following Mistral Small leads

GPT-4o mini: 61.9 (#239), Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkGPT-4o miniMistral Small
LiveBench Instruction Following56.8%63.7%
LMArena Instruction Following12581310
IFEval78.2%—

Long Context Mistral Small leads

GPT-4o mini: 39.1 (#186), Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkGPT-4o miniMistral Small
LMArena Longer Query12891327

Writing & Preference Mistral Small leads

GPT-4o mini: 39.5 (#248), Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkGPT-4o miniMistral Small
LMArena Text12861338
LMArena Creative Writing12681305
LMArena Multi-Turn12851344
LiveBench Language28.6%30.5%
Short-Story Creative Writing67.2%—
EQ-Bench Creative Writing873—
WildBench79.1%—

Frequently asked questions

Is GPT-4o mini better than Mistral Small?

Mistral Small is the stronger model overall, scoring 33.4 to 25.5 on the Noometry Index.

Which is cheaper, GPT-4o mini or Mistral Small?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; GPT-4o mini lists at $0.15 and $0.60.

Is GPT-4o mini or Mistral Small better for coding?

Mistral Small scores higher on coding benchmarks: 34.0 versus 22.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 128K.

How many benchmarks do GPT-4o mini and Mistral Small share?

34 benchmarks have published results for both models. GPT-4o mini has 60 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper