Model comparison

gpt-oss-20b vs Mistral Small 3.1

gpt-oss-20b and Mistral Small 3.1 score almost the same on the Noometry Index (32.5 vs 31.7), so choose on price, context window or the category you care about most.

Last verified . 26 shared benchmarks.

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 26 benchmarks with published results for both. gpt-oss-20b scores higher in 3 categories and Mistral Small 3.1 in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where gpt-oss-20b leads 39.4 to 14.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 65.3% for gpt-oss-20b and 3.9% for Mistral Small 3.1.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $0.35 / $0.56 for Mistral Small 3.1.
  • gpt-oss-20b accepts more context: 131K tokens versus 128K.

Side by side

gpt-oss-20b and Mistral Small 3.1 specifications
gpt-oss-20bMistral Small 3.1
ProviderOpenAIMistral AI
Noometry Index32.531.7
Released2025-08-052025-03-17
WeightsOpenOpen
Context window131K128K
Max output16K102K
Input $ / M tokens$0.018$0.35
Output $ / M tokens$0.09$0.56
Results tracked3428

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

gpt-oss-20b: 37.6 (#192), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
Benchmarkgpt-oss-20bMistral Small 3.1
LMArena Coding13061309
SciCode34.4%—
WeirdML40.9%—
ALE-Bench566.05—

Agentic & Tool Use Not comparable

gpt-oss-20b: 9.3 (#154), Mistral Small 3.1: —

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-20bMistral Small 3.1
Terminal-Bench3.4%—

Reasoning Too close to call

gpt-oss-20b: 19.3 (#261), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
Benchmarkgpt-oss-20bMistral Small 3.1
Chess Puzzles4%1%
LMArena Hard Prompts12741278
Epoch Capabilities Index137.82127.48
Kagi LLM Benchmark53.2%—
CritPt1.4%—
DTBench68%—
LMCA14.5%—

Math gpt-oss-20b leads

gpt-oss-20b: 39.4 (#103), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
Benchmarkgpt-oss-20bMistral Small 3.1
OTIS Mock AIME 2024-202565.3%3.9%
Omni-MATH56.5%24.8%
LMArena Math13171262

Knowledge gpt-oss-20b leads

gpt-oss-20b: 34.6 (#195), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
Benchmarkgpt-oss-20bMistral Small 3.1
GPQA Diamond60.8%41.9%
MMLU-Pro74%61%
GPQA (HELM)59.4%39.2%
LMArena Expert12581257

Multimodal Not comparable

gpt-oss-20b: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
Benchmarkgpt-oss-20bMistral Small 3.1
LMArena Vision—1136

Multilingual Too close to call

gpt-oss-20b: 42.2 (#197), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
Benchmarkgpt-oss-20bMistral Small 3.1
LMArena Non-English12681255
LMArena Chinese13141253
LMArena German12551266
LMArena Japanese12441208
LMArena Korean12361206
LMArena Russian12781263
LMArena Spanish12671283
LMArena French—1273

Instruction Following Mistral Small 3.1 leads

gpt-oss-20b: 61.8 (#240), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
Benchmarkgpt-oss-20bMistral Small 3.1
IFEval73.2%75%
LMArena Instruction Following12361264

Long Context Mistral Small 3.1 leads

gpt-oss-20b: 37.9 (#209), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
Benchmarkgpt-oss-20bMistral Small 3.1
LMArena Longer Query12501299

Writing & Preference Mistral Small 3.1 leads

gpt-oss-20b: 35.5 (#265), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
Benchmarkgpt-oss-20bMistral Small 3.1
LMArena Text12871277
LMArena Creative Writing12011253
EQ-Bench Creative Writing666761
WildBench73.7%78.8%
LMArena Multi-Turn12681270

Frequently asked questions

Is gpt-oss-20b better than Mistral Small 3.1?

gpt-oss-20b and Mistral Small 3.1 score almost the same on the Noometry Index (32.5 vs 31.7), so choose on price, context window or the category you care about most.

Which is cheaper, gpt-oss-20b or Mistral Small 3.1?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Mistral Small 3.1 lists at $0.35 and $0.56.

Is gpt-oss-20b or Mistral Small 3.1 better for coding?

They score almost the same on coding (37.6 vs 38.3); test both on your own repository before choosing.

Which has the bigger context window?

gpt-oss-20b does, with 131K tokens against 128K.

How many benchmarks do gpt-oss-20b and Mistral Small 3.1 share?

26 benchmarks have published results for both models. gpt-oss-20b has 34 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper