Model comparison

gpt-oss-20b vs Mistral Medium 3.1

gpt-oss-20b and Mistral Medium 3.1 score almost the same on the Noometry Index (32.5 vs 31.9), so choose on price, context window or the category you care about most.

Last verified . 1 shared benchmarks.

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Summary

  • They share 1 benchmark with published results for both. gpt-oss-20b scores higher in 1 category and Mistral Medium 3.1 in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium 3.1 leads 55.5 to 35.5.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $0.40 / $2 for Mistral Medium 3.1.
  • gpt-oss-20b has downloadable open weights; the other is API-only.

Side by side

gpt-oss-20b and Mistral Medium 3.1 specifications
gpt-oss-20bMistral Medium 3.1
ProviderOpenAIMistral AI
Noometry Index32.531.9
Released2025-08-05—
WeightsOpenProprietary
Context window131K131K
Max output16K105K
Input $ / M tokens$0.018$0.40
Output $ / M tokens$0.09$2
Results tracked343

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

gpt-oss-20b: 37.6 (#192), Mistral Medium 3.1: —

Coding benchmarks
Benchmarkgpt-oss-20bMistral Medium 3.1
SciCode34.4%—
WeirdML40.9%—
LMArena Coding1306—
ALE-Bench566.05—

Agentic & Tool Use Not comparable

gpt-oss-20b: 9.3 (#154), Mistral Medium 3.1: —

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-20bMistral Medium 3.1
Terminal-Bench3.4%—

Reasoning gpt-oss-20b leads

gpt-oss-20b: 19.3 (#261), Mistral Medium 3.1: 10.6 (#341)

Reasoning benchmarks
Benchmarkgpt-oss-20bMistral Medium 3.1
Kagi LLM Benchmark53.2%—
NYT Connections (extended)—6.5%
CritPt1.4%—
Chess Puzzles4%—
Thematic Generalization—20.3%
LMArena Hard Prompts1274—
DTBench68%—
LMCA14.5%—
Epoch Capabilities Index137.82—

Math Not comparable

gpt-oss-20b: 39.4 (#103), Mistral Medium 3.1: —

Math benchmarks
Benchmarkgpt-oss-20bMistral Medium 3.1
OTIS Mock AIME 2024-202565.3%—
Omni-MATH56.5%—
LMArena Math1317—

Knowledge Not comparable

gpt-oss-20b: 34.6 (#195), Mistral Medium 3.1: —

Knowledge benchmarks
Benchmarkgpt-oss-20bMistral Medium 3.1
GPQA Diamond60.8%—
MMLU-Pro74%—
GPQA (HELM)59.4%—
LMArena Expert1258—

Multilingual Not comparable

gpt-oss-20b: 42.2 (#197), Mistral Medium 3.1: —

Multilingual benchmarks
Benchmarkgpt-oss-20bMistral Medium 3.1
LMArena Non-English1268—
LMArena Chinese1314—
LMArena German1255—
LMArena Japanese1244—
LMArena Korean1236—
LMArena Russian1278—
LMArena Spanish1267—

Instruction Following Not comparable

gpt-oss-20b: 61.8 (#240), Mistral Medium 3.1: —

Instruction Following benchmarks
Benchmarkgpt-oss-20bMistral Medium 3.1
IFEval73.2%—
LMArena Instruction Following1236—

Long Context Not comparable

gpt-oss-20b: 37.9 (#209), Mistral Medium 3.1: —

Long Context benchmarks
Benchmarkgpt-oss-20bMistral Medium 3.1
LMArena Longer Query1250—

Writing & Preference Mistral Medium 3.1 leads

gpt-oss-20b: 35.5 (#265), Mistral Medium 3.1: 55.5 (#145)

Writing & Preference benchmarks
Benchmarkgpt-oss-20bMistral Medium 3.1
EQ-Bench Creative Writing6661476
LMArena Text1287—
LMArena Creative Writing1201—
WildBench73.7%—
LMArena Multi-Turn1268—

Frequently asked questions

Is gpt-oss-20b better than Mistral Medium 3.1?

gpt-oss-20b and Mistral Medium 3.1 score almost the same on the Noometry Index (32.5 vs 31.9), so choose on price, context window or the category you care about most.

Which is cheaper, gpt-oss-20b or Mistral Medium 3.1?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Mistral Medium 3.1 lists at $0.40 and $2.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do gpt-oss-20b and Mistral Medium 3.1 share?

1 benchmark has published results for both models. gpt-oss-20b has 34 scored results on Noometry and Mistral Medium 3.1 has 3.

Related comparisons

Go deeper