Model comparison

GPT-3.5-turbo vs Mistral

Mistral is the stronger model overall, scoring 29.9 to 23.2 on the Noometry Index.

Last verified . 17 shared benchmarks.

GPT-3.5-turbo OpenAI

23.2

Rank #350 Confirmed

Mistral Mistral AI

29.9

Rank #303 Confirmed

Summary

  • They share 17 benchmarks with published results for both. GPT-3.5-turbo scores higher in 1 category and Mistral in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral leads 22.3 to 6.3.

Side by side

GPT-3.5-turbo and Mistral specifications
GPT-3.5-turboMistral
ProviderOpenAIMistral AI
Noometry Index23.229.9
Released2023-03-01—
WeightsProprietaryProprietary
Context window16K—
Max output4K—
Input $ / M tokens$0.50—
Output $ / M tokens$1.50—
Results tracked4422

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

GPT-3.5-turbo: 23.9 (#331), Mistral: 33.8 (#250)

Coding benchmarks
BenchmarkGPT-3.5-turboMistral
LMArena Coding11361162
WeirdML3.5%—
BigCodeBench Instruct39.1%—
BigCodeBench Complete50.6%—
HumanEval+70.7%—
MBPP+69.7%—

Agentic & Tool Use Not comparable

GPT-3.5-turbo: —, Mistral: —

Agentic & Tool Use benchmarks
BenchmarkGPT-3.5-turboMistral
METR Time Horizons21.5%—

Reasoning Mistral leads

GPT-3.5-turbo: 13.8 (#332), Mistral: 22.2 (#200)

Reasoning benchmarks
BenchmarkGPT-3.5-turboMistral
LMArena Hard Prompts11081149
Chess Puzzles0%—
Mystery Game Puzzles3%—
DTBench48.5%—
LMCA9.7%—
Adversarial NLI58.1%—
BIG-Bench Hard61.6%—
CommonsenseQA 2.057%—
Epoch Capabilities Index118.55—
ForecastBench50.4—
WinoGrande81.6%—

Math Mistral leads

GPT-3.5-turbo: 6.3 (#327), Mistral: 22.3 (#278)

Math benchmarks
BenchmarkGPT-3.5-turboMistral
LMArena Math11421180
FrontierMath (Tiers 1-3)0%—
OTIS Mock AIME 2024-20252.2%—
Omni-MATH—7.2%
MATH Level 515.9%—
GSM8K57.8%—

Knowledge Mistral leads

GPT-3.5-turbo: 10.0 (#303), Mistral: 16.6 (#288)

Knowledge benchmarks
BenchmarkGPT-3.5-turboMistral
LMArena Expert10701125
GPQA Diamond28%—
MMLU-Pro—27.7%
GPQA (HELM)—30.3%
ARC (AI2) Challenge87.4%—
BoolQ87%—
MMLU71.4%—
OpenBookQA86%—
TriviaQA85.8%—

Multilingual Mistral leads

GPT-3.5-turbo: 31.5 (#258), Mistral: 32.8 (#254)

Multilingual benchmarks
BenchmarkGPT-3.5-turboMistral
LMArena Non-English11081129
LMArena Chinese10751109
LMArena French11181180
LMArena German10901155
LMArena Japanese10431013
LMArena Korean10191032
LMArena Russian11231168
LMArena Spanish11211143

Instruction Following GPT-3.5-turbo leads

GPT-3.5-turbo: 57.9 (#262), Mistral: 52.6 (#288)

Instruction Following benchmarks
BenchmarkGPT-3.5-turboMistral
LMArena Instruction Following11191152
IFEval—56.8%

Long Context Too close to call

GPT-3.5-turbo: 34.0 (#254), Mistral: 35.0 (#245)

Long Context benchmarks
BenchmarkGPT-3.5-turboMistral
LMArena Longer Query11211153

Writing & Preference Mistral leads

GPT-3.5-turbo: 25.3 (#305), Mistral: 37.0 (#260)

Writing & Preference benchmarks
BenchmarkGPT-3.5-turboMistral
LMArena Text11251165
LMArena Creative Writing10921158
LMArena Multi-Turn11171147
EQ-Bench Creative Writing451—
WildBench—66%

Frequently asked questions

Is GPT-3.5-turbo better than Mistral?

Mistral is the stronger model overall, scoring 29.9 to 23.2 on the Noometry Index.

Is GPT-3.5-turbo or Mistral better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 23.9 in the Noometry coding category.

How many benchmarks do GPT-3.5-turbo and Mistral share?

17 benchmarks have published results for both models. GPT-3.5-turbo has 44 scored results on Noometry and Mistral has 22.

Related comparisons

Go deeper