Model comparison

gpt-oss-120b vs Magistral Medium

gpt-oss-120b is the stronger model overall, scoring 36.3 to 35.2 on the Noometry Index.

Last verified . 20 shared benchmarks.

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 20 benchmarks with published results for both. gpt-oss-120b scores higher in 6 categories and Magistral Medium in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where gpt-oss-120b leads 52.5 to 35.1.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 58.6% for gpt-oss-120b and 16.2% for Magistral Medium.
  • gpt-oss-120b is cheaper at $0.037 / $0.17 per million input/output tokens, against $2 / $5 for Magistral Medium.
  • Magistral Medium accepts more context: 262K tokens versus 131K.

Side by side

gpt-oss-120b and Magistral Medium specifications
gpt-oss-120bMagistral Medium
ProviderOpenAIMistral AI
Noometry Index36.335.2
Released2025-08-052025-03-17
WeightsOpenOpen
Context window131K262K
Max output41K16K
Input $ / M tokens$0.037$2
Output $ / M tokens$0.17$5
Results tracked4822

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

gpt-oss-120b: 33.5 (#256), Magistral Medium: 39.1 (#161)

Coding benchmarks
Benchmarkgpt-oss-120bMagistral Medium
SciCode36%39.2%
LMArena Coding13801319
SWE-bench Verified (bash only)26%—
Aider Polyglot41.8%—
WeirdML48.2%—
ALE-Bench575.62—
AlgoTune1.41—

Agentic & Tool Use Not comparable

gpt-oss-120b: 12.2 (#153), Magistral Medium: —

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-120bMagistral Medium
Terminal-Bench18.7%—
APEX-Agents4.4%—
METR Time Horizons56.6%—
Vending-Bench 2-21.53—

Reasoning gpt-oss-120b leads

gpt-oss-120b: 20.0 (#245), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
Benchmarkgpt-oss-120bMagistral Medium
Kagi LLM Benchmark58.6%16.2%
CritPt1.1%0.3%
LMArena Hard Prompts13641267
ARC-AGI-2—0%
SimpleBench22.1%—
ARC-AGI-1—6.1%
Chess Puzzles20%—
Mystery Game Puzzles2%—
DTBench76.3%—
LMCA22.1%—
Surface Evolver Bench25%—
Epoch Capabilities Index139.93—

Math gpt-oss-120b leads

gpt-oss-120b: 52.5 (#50), Magistral Medium: 35.1 (#189)

Math benchmarks
Benchmarkgpt-oss-120bMagistral Medium
LMArena Math13891250
OTIS Mock AIME 2024-202588.9%—
Omni-MATH68.8%—

Knowledge gpt-oss-120b leads

gpt-oss-120b: 42.4 (#96), Magistral Medium: 33.5 (#202)

Knowledge benchmarks
Benchmarkgpt-oss-120bMagistral Medium
LMArena Expert13561223
GPQA Diamond75.8%—
MMLU-Pro79.5%—
Confabulations15.7%—
Vectara Hallucination Rate14.2%—
GPQA (HELM)68.4%—

Multilingual gpt-oss-120b leads

gpt-oss-120b: 48.0 (#147), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
Benchmarkgpt-oss-120bMagistral Medium
LMArena Non-English13511232
LMArena Chinese13851227
LMArena French13691267
LMArena German13531248
LMArena Japanese13311175
LMArena Korean12821125
LMArena Russian13431224
LMArena Spanish13891271

Instruction Following gpt-oss-120b leads

gpt-oss-120b: 69.3 (#173), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
Benchmarkgpt-oss-120bMagistral Medium
LMArena Instruction Following13181254
IFEval83.6%—

Long Context Magistral Medium leads

gpt-oss-120b: 31.4 (#278), Magistral Medium: 39.3 (#183)

Long Context benchmarks
Benchmarkgpt-oss-120bMagistral Medium
LMArena Longer Query13191295
Fiction.LiveBench44.4%—

Writing & Preference Too close to call

gpt-oss-120b: 46.5 (#217), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
Benchmarkgpt-oss-120bMagistral Medium
LMArena Text13651255
LMArena Creative Writing12751245
LMArena Multi-Turn13401275
Short-Story Creative Writing77.1%—
EQ-Bench Creative Writing961—
WildBench84.5%—

Frequently asked questions

Is gpt-oss-120b better than Magistral Medium?

gpt-oss-120b is the stronger model overall, scoring 36.3 to 35.2 on the Noometry Index.

Which is cheaper, gpt-oss-120b or Magistral Medium?

gpt-oss-120b is cheaper. It lists at $0.037 per million input tokens and $0.17 per million output tokens; Magistral Medium lists at $2 and $5.

Is gpt-oss-120b or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 131K.

How many benchmarks do gpt-oss-120b and Magistral Medium share?

20 benchmarks have published results for both models. gpt-oss-120b has 48 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper