Model comparison

gpt-oss-20b vs Mistral Large

gpt-oss-20b and Mistral Large score almost the same on the Noometry Index (32.5 vs 31.9), so choose on price, context window or the category you care about most.

Last verified . 30 shared benchmarks.

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 30 benchmarks with published results for both. gpt-oss-20b scores higher in 5 categories and Mistral Large in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where gpt-oss-20b leads 39.4 to 18.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 65.3% for gpt-oss-20b and 8.5% for Mistral Large.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $2 / $6 for Mistral Large.

Side by side

gpt-oss-20b and Mistral Large specifications
gpt-oss-20bMistral Large
ProviderOpenAIMistral AI
Noometry Index32.531.9
Released2025-08-052024-02-26
WeightsOpenOpen
Context window131K131K
Max output16K16K
Input $ / M tokens$0.018$2
Output $ / M tokens$0.09$6
Results tracked3451

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding gpt-oss-20b leads

gpt-oss-20b: 37.6 (#192), Mistral Large: 34.3 (#240)

Coding benchmarks
Benchmarkgpt-oss-20bMistral Large
SciCode34.4%36.2%
LMArena Coding13061277
ALE-Bench566.05264.7
WeirdML40.9%—
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Mistral Large leads

gpt-oss-20b: 9.3 (#154), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-20bMistral Large
Terminal-Bench3.4%—
Berkeley Function Calling Leaderboard—38.4%

Reasoning gpt-oss-20b leads

gpt-oss-20b: 19.3 (#261), Mistral Large: 15.8 (#310)

Reasoning benchmarks
Benchmarkgpt-oss-20bMistral Large
CritPt1.4%0%
LMArena Hard Prompts12741257
DTBench68%65.1%
LMCA14.5%16.7%
Epoch Capabilities Index137.82128.52
SimpleBench—22.5%
Kagi LLM Benchmark53.2%—
Chess Puzzles4%—
LiveBench Reasoning—43.5%
LiveBench Data Analysis—50.1%
ForecastBench—57.1
LiveBench—48.4%

Math gpt-oss-20b leads

gpt-oss-20b: 39.4 (#103), Mistral Large: 18.2 (#291)

Math benchmarks
Benchmarkgpt-oss-20bMistral Large
OTIS Mock AIME 2024-202565.3%8.5%
Omni-MATH56.5%28.1%
LMArena Math13171262
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge gpt-oss-20b leads

gpt-oss-20b: 34.6 (#195), Mistral Large: 30.1 (#230)

Knowledge benchmarks
Benchmarkgpt-oss-20bMistral Large
GPQA Diamond60.8%51.3%
MMLU-Pro74%59.9%
GPQA (HELM)59.4%43.5%
LMArena Expert12581232
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
MMLU—80%

Multilingual gpt-oss-20b leads

gpt-oss-20b: 42.2 (#197), Mistral Large: 40.0 (#219)

Multilingual benchmarks
Benchmarkgpt-oss-20bMistral Large
LMArena Non-English12681237
LMArena Chinese13141240
LMArena German12551254
LMArena Japanese12441188
LMArena Korean12361202
LMArena Russian12781257
LMArena Spanish12671268
LMArena French—1325

Instruction Following Mistral Large leads

gpt-oss-20b: 61.8 (#240), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
Benchmarkgpt-oss-20bMistral Large
IFEval73.2%87.7%
LMArena Instruction Following12361249
LiveBench Instruction Following—67.9%

Long Context Too close to call

gpt-oss-20b: 37.9 (#209), Mistral Large: 38.3 (#199)

Long Context benchmarks
Benchmarkgpt-oss-20bMistral Large
LMArena Longer Query12501261

Writing & Preference Mistral Large leads

gpt-oss-20b: 35.5 (#265), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
Benchmarkgpt-oss-20bMistral Large
LMArena Text12871266
LMArena Creative Writing12011243
EQ-Bench Creative Writing666985
WildBench73.7%80.1%
LMArena Multi-Turn12681260
Short-Story Creative Writing—69%
LiveBench Language—39.4%

Frequently asked questions

Is gpt-oss-20b better than Mistral Large?

gpt-oss-20b and Mistral Large score almost the same on the Noometry Index (32.5 vs 31.9), so choose on price, context window or the category you care about most.

Which is cheaper, gpt-oss-20b or Mistral Large?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Mistral Large lists at $2 and $6.

Is gpt-oss-20b or Mistral Large better for coding?

gpt-oss-20b scores higher on coding benchmarks: 37.6 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do gpt-oss-20b and Mistral Large share?

30 benchmarks have published results for both models. gpt-oss-20b has 34 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper