Model comparison

gpt-oss-20b vs Mixtral 8x7B

gpt-oss-20b is the stronger model overall, scoring 32.5 to 27.1 on the Noometry Index.

Last verified . 24 shared benchmarks.

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Summary

  • They share 24 benchmarks with published results for both. gpt-oss-20b scores higher in 8 categories and Mixtral 8x7B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where gpt-oss-20b leads 34.6 to 11.0.
  • The biggest single-benchmark swing is Omni-MATH: 56.5% for gpt-oss-20b and 10.5% for Mixtral 8x7B.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $0.70 / $0.70 for Mixtral 8x7B.
  • gpt-oss-20b accepts more context: 131K tokens versus 32K.

Side by side

gpt-oss-20b and Mixtral 8x7B specifications
gpt-oss-20bMixtral 8x7B
ProviderOpenAIMistral AI
Noometry Index32.527.1
Released2025-08-052023-12-11
WeightsOpenOpen
Context window131K32K
Max output16K32K
Input $ / M tokens$0.018$0.70
Output $ / M tokens$0.09$0.70
Results tracked3438

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding gpt-oss-20b leads

gpt-oss-20b: 37.6 (#192), Mixtral 8x7B: 32.8 (#269)

Coding benchmarks
Benchmarkgpt-oss-20bMixtral 8x7B
LMArena Coding13061126
SciCode34.4%—
WeirdML40.9%—
ALE-Bench566.05—
HumanEval+—39.6%
MBPP+—49.7%

Agentic & Tool Use Not comparable

gpt-oss-20b: 9.3 (#154), Mixtral 8x7B: —

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-20bMixtral 8x7B
Terminal-Bench3.4%—

Reasoning gpt-oss-20b leads

gpt-oss-20b: 19.3 (#261), Mixtral 8x7B: 18.2 (#285)

Reasoning benchmarks
Benchmarkgpt-oss-20bMixtral 8x7B
LMArena Hard Prompts12741115
DTBench68%49.6%
Epoch Capabilities Index137.82118.47
Kagi LLM Benchmark53.2%—
CritPt1.4%—
Chess Puzzles4%—
LMCA14.5%—
Adversarial NLI—55.2%
ForecastBench—56.3
HellaSwag—86.7%
PIQA—83.6%
WinoGrande—77.2%

Math gpt-oss-20b leads

gpt-oss-20b: 39.4 (#103), Mixtral 8x7B: 18.8 (#289)

Math benchmarks
Benchmarkgpt-oss-20bMixtral 8x7B
Omni-MATH56.5%10.5%
LMArena Math13171147
OTIS Mock AIME 2024-202565.3%—
MATH Level 5—10%
GSM8K—74.4%

Knowledge gpt-oss-20b leads

gpt-oss-20b: 34.6 (#195), Mixtral 8x7B: 11.0 (#301)

Knowledge benchmarks
Benchmarkgpt-oss-20bMixtral 8x7B
GPQA Diamond60.8%30.6%
MMLU-Pro74%33.5%
GPQA (HELM)59.4%29.6%
LMArena Expert12581088
ARC (AI2) Challenge—87.3%
MMLU—70.6%
OpenBookQA—85.8%
TriviaQA—82.2%

Multilingual gpt-oss-20b leads

gpt-oss-20b: 42.2 (#197), Mixtral 8x7B: 29.6 (#266)

Multilingual benchmarks
Benchmarkgpt-oss-20bMixtral 8x7B
LMArena Non-English12681077
LMArena Chinese13141055
LMArena German12551114
LMArena Japanese1244931
LMArena Korean1236968
LMArena Russian12781090
LMArena Spanish12671111
LMArena French—1166

Instruction Following gpt-oss-20b leads

gpt-oss-20b: 61.8 (#240), Mixtral 8x7B: 51.0 (#297)

Instruction Following benchmarks
Benchmarkgpt-oss-20bMixtral 8x7B
IFEval73.2%57.5%
LMArena Instruction Following12361109

Long Context gpt-oss-20b leads

gpt-oss-20b: 37.9 (#209), Mixtral 8x7B: 33.4 (#260)

Long Context benchmarks
Benchmarkgpt-oss-20bMixtral 8x7B
LMArena Longer Query12501103

Writing & Preference gpt-oss-20b leads

gpt-oss-20b: 35.5 (#265), Mixtral 8x7B: 34.2 (#270)

Writing & Preference benchmarks
Benchmarkgpt-oss-20bMixtral 8x7B
LMArena Text12871132
LMArena Creative Writing12011109
WildBench73.7%67.3%
LMArena Multi-Turn12681115
EQ-Bench Creative Writing666—

Frequently asked questions

Is gpt-oss-20b better than Mixtral 8x7B?

gpt-oss-20b is the stronger model overall, scoring 32.5 to 27.1 on the Noometry Index.

Which is cheaper, gpt-oss-20b or Mixtral 8x7B?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Mixtral 8x7B lists at $0.70 and $0.70.

Is gpt-oss-20b or Mixtral 8x7B better for coding?

gpt-oss-20b scores higher on coding benchmarks: 37.6 versus 32.8 in the Noometry coding category.

Which has the bigger context window?

gpt-oss-20b does, with 131K tokens against 32K.

How many benchmarks do gpt-oss-20b and Mixtral 8x7B share?

24 benchmarks have published results for both models. gpt-oss-20b has 34 scored results on Noometry and Mixtral 8x7B has 38.

Related comparisons

Go deeper