Model comparison

GPT-5.3 Chat vs Magistral Medium

GPT-5.3 Chat is the stronger model overall, scoring 42.8 to 35.2 on the Noometry Index. Magistral Medium costs 1.8× less per token, which makes it the better buy when GPT-5.3 Chat's lead doesn't matter for your workload.

Last verified . 17 shared benchmarks.

GPT-5.3 Chat OpenAI

42.8

Rank #109 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 17 benchmarks with published results for both. GPT-5.3 Chat scores higher in 8 categories and Magistral Medium in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-5.3 Chat leads 28.5 to 8.6.
  • Magistral Medium is cheaper at $2 / $5 per million input/output tokens, against $1.75 / $14 for GPT-5.3 Chat.
  • Magistral Medium accepts more context: 262K tokens versus 128K.
  • Magistral Medium has downloadable open weights; the other is API-only.

Side by side

GPT-5.3 Chat and Magistral Medium specifications
GPT-5.3 ChatMagistral Medium
ProviderOpenAIMistral AI
Noometry Index42.835.2
Released2026-03-032025-03-17
WeightsProprietaryOpen
Context window128K262K
Max output16K16K
Input $ / M tokens$1.75$2
Output $ / M tokens$14$5
Results tracked1822

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.3 Chat leads

GPT-5.3 Chat: 41.4 (#124), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkGPT-5.3 ChatMagistral Medium
LMArena Coding14081319
SciCode—39.2%

Reasoning GPT-5.3 Chat leads

GPT-5.3 Chat: 28.5 (#102), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkGPT-5.3 ChatMagistral Medium
LMArena Hard Prompts13991267
ARC-AGI-2—0%
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
CritPt—0.3%

Math GPT-5.3 Chat leads

GPT-5.3 Chat: 38.2 (#142), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkGPT-5.3 ChatMagistral Medium
LMArena Math13891250

Knowledge GPT-5.3 Chat leads

GPT-5.3 Chat: 38.8 (#140), Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkGPT-5.3 ChatMagistral Medium
LMArena Expert13971223

Multilingual GPT-5.3 Chat leads

GPT-5.3 Chat: 50.3 (#124), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkGPT-5.3 ChatMagistral Medium
LMArena Non-English13821232
LMArena Chinese14321227
LMArena French13971267
LMArena German13841248
LMArena Japanese13521175
LMArena Korean13461125
LMArena Russian14001224
LMArena Spanish13711271

Instruction Following GPT-5.3 Chat leads

GPT-5.3 Chat: 72.8 (#129), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkGPT-5.3 ChatMagistral Medium
LMArena Instruction Following13781254

Long Context GPT-5.3 Chat leads

GPT-5.3 Chat: 42.6 (#120), Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkGPT-5.3 ChatMagistral Medium
LMArena Longer Query13961295

Writing & Preference GPT-5.3 Chat leads

GPT-5.3 Chat: 63.1 (#68), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkGPT-5.3 ChatMagistral Medium
LMArena Text13891255
LMArena Creative Writing13551245
LMArena Multi-Turn14121275
EQ-Bench Creative Writing1690—

Frequently asked questions

Is GPT-5.3 Chat better than Magistral Medium?

GPT-5.3 Chat is the stronger model overall, scoring 42.8 to 35.2 on the Noometry Index. Magistral Medium costs 1.8× less per token, which makes it the better buy when GPT-5.3 Chat's lead doesn't matter for your workload.

Which is cheaper, GPT-5.3 Chat or Magistral Medium?

Magistral Medium is cheaper. It lists at $2 per million input tokens and $5 per million output tokens; GPT-5.3 Chat lists at $1.75 and $14.

Is GPT-5.3 Chat or Magistral Medium better for coding?

GPT-5.3 Chat scores higher on coding benchmarks: 41.4 versus 39.1 in the Noometry coding category.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 128K.

How many benchmarks do GPT-5.3 Chat and Magistral Medium share?

17 benchmarks have published results for both models. GPT-5.3 Chat has 18 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper