Model comparison
gpt-oss-20b vs Mistral Medium 3.1
gpt-oss-20b and Mistral Medium 3.1 score almost the same on the Noometry Index (32.5 vs 31.9), so choose on price, context window or the category you care about most.
Last verified . 1 shared benchmarks.
Summary
- They share 1 benchmark with published results for both. gpt-oss-20b scores higher in 1 category and Mistral Medium 3.1 in 1 category; 2 gaps are clear of the uncertainty.
- The widest gap is in writing & preference, where Mistral Medium 3.1 leads 55.5 to 35.5.
- gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $0.40 / $2 for Mistral Medium 3.1.
- gpt-oss-20b has downloadable open weights; the other is API-only.
Side by side
| gpt-oss-20b | Mistral Medium 3.1 | |
|---|---|---|
| Provider | OpenAI | Mistral AI |
| Noometry Index | 32.5 | 31.9 |
| Released | 2025-08-05 | — |
| Weights | Open | Proprietary |
| Context window | 131K | 131K |
| Max output | 16K | 105K |
| Input $ / M tokens | $0.018 | $0.40 |
| Output $ / M tokens | $0.09 | $2 |
| Results tracked | 34 | 3 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
gpt-oss-20b: 37.6 (#192), Mistral Medium 3.1: —
| Benchmark | gpt-oss-20b | Mistral Medium 3.1 |
|---|---|---|
| SciCode | 34.4% | — |
| WeirdML | 40.9% | — |
| LMArena Coding | 1306 | — |
| ALE-Bench | 566.05 | — |
Agentic & Tool Use Not comparable
gpt-oss-20b: 9.3 (#154), Mistral Medium 3.1: —
| Benchmark | gpt-oss-20b | Mistral Medium 3.1 |
|---|---|---|
| Terminal-Bench | 3.4% | — |
Reasoning gpt-oss-20b leads
gpt-oss-20b: 19.3 (#261), Mistral Medium 3.1: 10.6 (#341)
| Benchmark | gpt-oss-20b | Mistral Medium 3.1 |
|---|---|---|
| Kagi LLM Benchmark | 53.2% | — |
| NYT Connections (extended) | — | 6.5% |
| CritPt | 1.4% | — |
| Chess Puzzles | 4% | — |
| Thematic Generalization | — | 20.3% |
| LMArena Hard Prompts | 1274 | — |
| DTBench | 68% | — |
| LMCA | 14.5% | — |
| Epoch Capabilities Index | 137.82 | — |
Math Not comparable
gpt-oss-20b: 39.4 (#103), Mistral Medium 3.1: —
| Benchmark | gpt-oss-20b | Mistral Medium 3.1 |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 65.3% | — |
| Omni-MATH | 56.5% | — |
| LMArena Math | 1317 | — |
Knowledge Not comparable
gpt-oss-20b: 34.6 (#195), Mistral Medium 3.1: —
| Benchmark | gpt-oss-20b | Mistral Medium 3.1 |
|---|---|---|
| GPQA Diamond | 60.8% | — |
| MMLU-Pro | 74% | — |
| GPQA (HELM) | 59.4% | — |
| LMArena Expert | 1258 | — |
Multilingual Not comparable
gpt-oss-20b: 42.2 (#197), Mistral Medium 3.1: —
| Benchmark | gpt-oss-20b | Mistral Medium 3.1 |
|---|---|---|
| LMArena Non-English | 1268 | — |
| LMArena Chinese | 1314 | — |
| LMArena German | 1255 | — |
| LMArena Japanese | 1244 | — |
| LMArena Korean | 1236 | — |
| LMArena Russian | 1278 | — |
| LMArena Spanish | 1267 | — |
Instruction Following Not comparable
gpt-oss-20b: 61.8 (#240), Mistral Medium 3.1: —
| Benchmark | gpt-oss-20b | Mistral Medium 3.1 |
|---|---|---|
| IFEval | 73.2% | — |
| LMArena Instruction Following | 1236 | — |
Long Context Not comparable
gpt-oss-20b: 37.9 (#209), Mistral Medium 3.1: —
| Benchmark | gpt-oss-20b | Mistral Medium 3.1 |
|---|---|---|
| LMArena Longer Query | 1250 | — |
Writing & Preference Mistral Medium 3.1 leads
gpt-oss-20b: 35.5 (#265), Mistral Medium 3.1: 55.5 (#145)
| Benchmark | gpt-oss-20b | Mistral Medium 3.1 |
|---|---|---|
| EQ-Bench Creative Writing | 666 | 1476 |
| LMArena Text | 1287 | — |
| LMArena Creative Writing | 1201 | — |
| WildBench | 73.7% | — |
| LMArena Multi-Turn | 1268 | — |
Frequently asked questions
Is gpt-oss-20b better than Mistral Medium 3.1?
gpt-oss-20b and Mistral Medium 3.1 score almost the same on the Noometry Index (32.5 vs 31.9), so choose on price, context window or the category you care about most.
Which is cheaper, gpt-oss-20b or Mistral Medium 3.1?
gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Mistral Medium 3.1 lists at $0.40 and $2.
Which has the bigger context window?
Both accept 131K tokens.
How many benchmarks do gpt-oss-20b and Mistral Medium 3.1 share?
1 benchmark has published results for both models. gpt-oss-20b has 34 scored results on Noometry and Mistral Medium 3.1 has 3.