Model comparison

gpt-oss-20b vs Mistral Small

gpt-oss-20b and Mistral Small score almost the same on the Noometry Index (32.5 vs 33.4), so choose on price, context window or the category you care about most.

Last verified . 24 shared benchmarks.

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 24 benchmarks with published results for both. gpt-oss-20b scores higher in 3 categories and Mistral Small in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where gpt-oss-20b leads 39.4 to 16.4.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 65.3% for gpt-oss-20b and 5.8% for Mistral Small.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $0.15 / $0.60 for Mistral Small.
  • Mistral Small accepts more context: 262K tokens versus 131K.

Side by side

gpt-oss-20b and Mistral Small specifications
gpt-oss-20bMistral Small
ProviderOpenAIMistral AI
Noometry Index32.533.4
Released2025-08-052024-02-26
WeightsOpenOpen
Context window131K262K
Max output16K256K
Input $ / M tokens$0.018$0.15
Output $ / M tokens$0.09$0.60
Results tracked3439

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding gpt-oss-20b leads

gpt-oss-20b: 37.6 (#192), Mistral Small: 34.0 (#247)

Coding benchmarks
Benchmarkgpt-oss-20bMistral Small
SciCode34.4%26.5%
LMArena Coding13061362
ALE-Bench566.05497.62
WeirdML40.9%—
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
BigCodeBench Complete—46.6%

Agentic & Tool Use Mistral Small leads

gpt-oss-20b: 9.3 (#154), Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-20bMistral Small
Terminal-Bench3.4%—
Berkeley Function Calling Leaderboard—37.1%

Reasoning Too close to call

gpt-oss-20b: 19.3 (#261), Mistral Small: 19.8 (#250)

Reasoning benchmarks
Benchmarkgpt-oss-20bMistral Small
Kagi LLM Benchmark53.2%37.8%
CritPt1.4%0%
LMArena Hard Prompts12741335
DTBench68%70.9%
LMCA14.5%20.6%
Chess Puzzles4%—
LiveBench Reasoning—44.8%
LiveBench Data Analysis—53.7%
Epoch Capabilities Index137.82—
LiveBench—44%

Math gpt-oss-20b leads

gpt-oss-20b: 39.4 (#103), Mistral Small: 16.4 (#293)

Math benchmarks
Benchmarkgpt-oss-20bMistral Small
OTIS Mock AIME 2024-202565.3%5.8%
LMArena Math13171341
Omni-MATH56.5%—
LiveBench Math—39.9%
MATH Level 5—46.8%

Knowledge gpt-oss-20b leads

gpt-oss-20b: 34.6 (#195), Mistral Small: 31.0 (#222)

Knowledge benchmarks
Benchmarkgpt-oss-20bMistral Small
GPQA Diamond60.8%47.5%
LMArena Expert12581291
MMLU-Pro74%—
Vectara Hallucination Rate—5.1%
GPQA (HELM)59.4%—
MMLU—68.7%

Multimodal Not comparable

gpt-oss-20b: —, Mistral Small: 33.5 (#96)

Multimodal benchmarks
Benchmarkgpt-oss-20bMistral Small
LMArena Vision—1142

Multilingual Mistral Small leads

gpt-oss-20b: 42.2 (#197), Mistral Small: 45.5 (#169)

Multilingual benchmarks
Benchmarkgpt-oss-20bMistral Small
LMArena Non-English12681315
LMArena Chinese13141340
LMArena German12551340
LMArena Japanese12441275
LMArena Korean12361259
LMArena Russian12781324
LMArena Spanish12671346
LMArena French—1337

Instruction Following Mistral Small leads

gpt-oss-20b: 61.8 (#240), Mistral Small: 66.4 (#209)

Instruction Following benchmarks
Benchmarkgpt-oss-20bMistral Small
LMArena Instruction Following12361310
LiveBench Instruction Following—63.7%
IFEval73.2%—

Long Context Mistral Small leads

gpt-oss-20b: 37.9 (#209), Mistral Small: 40.4 (#156)

Long Context benchmarks
Benchmarkgpt-oss-20bMistral Small
LMArena Longer Query12501327

Writing & Preference Mistral Small leads

gpt-oss-20b: 35.5 (#265), Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
Benchmarkgpt-oss-20bMistral Small
LMArena Text12871338
LMArena Creative Writing12011305
LMArena Multi-Turn12681344
EQ-Bench Creative Writing666—
WildBench73.7%—
LiveBench Language—30.5%

Frequently asked questions

Is gpt-oss-20b better than Mistral Small?

gpt-oss-20b and Mistral Small score almost the same on the Noometry Index (32.5 vs 33.4), so choose on price, context window or the category you care about most.

Which is cheaper, gpt-oss-20b or Mistral Small?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Mistral Small lists at $0.15 and $0.60.

Is gpt-oss-20b or Mistral Small better for coding?

gpt-oss-20b scores higher on coding benchmarks: 37.6 versus 34.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 131K.

How many benchmarks do gpt-oss-20b and Mistral Small share?

24 benchmarks have published results for both models. gpt-oss-20b has 34 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper