Model comparison

gpt-oss-20b vs Pixtral Large

gpt-oss-20b and Pixtral Large score almost the same on the Noometry Index (32.5 vs 32.2), so choose on price, context window or the category you care about most.

Last verified . 1 shared benchmarks.

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Summary

  • They share 1 benchmark with published results for both. gpt-oss-20b scores higher in 1 category and Pixtral Large in 1 category; 2 gaps are clear of the uncertainty.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $2 / $6 for Pixtral Large.
  • gpt-oss-20b accepts more context: 131K tokens versus 128K.

Side by side

gpt-oss-20b and Pixtral Large specifications
gpt-oss-20bPixtral Large
ProviderOpenAIMistral AI
Noometry Index32.532.2
Released2025-08-052024-11-01
WeightsOpenOpen
Context window131K128K
Max output16K128K
Input $ / M tokens$0.018$2
Output $ / M tokens$0.09$6
Results tracked343

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

gpt-oss-20b: 37.6 (#192), Pixtral Large: —

Coding benchmarks
Benchmarkgpt-oss-20bPixtral Large
SciCode34.4%—
WeirdML40.9%—
LMArena Coding1306—
ALE-Bench566.05—

Agentic & Tool Use Not comparable

gpt-oss-20b: 9.3 (#154), Pixtral Large: —

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-20bPixtral Large
Terminal-Bench3.4%—

Reasoning Pixtral Large leads

gpt-oss-20b: 19.3 (#261), Pixtral Large: 21.7 (#218)

Reasoning benchmarks
Benchmarkgpt-oss-20bPixtral Large
Kagi LLM Benchmark53.2%—
CritPt1.4%—
Chess Puzzles4%—
EnigmaEval—0.8%
LMArena Hard Prompts1274—
DTBench68%—
LMCA14.5%—
Epoch Capabilities Index137.82—

Math Not comparable

gpt-oss-20b: 39.4 (#103), Pixtral Large: —

Math benchmarks
Benchmarkgpt-oss-20bPixtral Large
OTIS Mock AIME 2024-202565.3%—
Omni-MATH56.5%—
LMArena Math1317—

Knowledge Not comparable

gpt-oss-20b: 34.6 (#195), Pixtral Large: —

Knowledge benchmarks
Benchmarkgpt-oss-20bPixtral Large
GPQA Diamond60.8%—
MMLU-Pro74%—
GPQA (HELM)59.4%—
LMArena Expert1258—

Multimodal Not comparable

gpt-oss-20b: —, Pixtral Large: 30.6 (#111)

Multimodal benchmarks
Benchmarkgpt-oss-20bPixtral Large
LMArena Vision—1089

Multilingual Not comparable

gpt-oss-20b: 42.2 (#197), Pixtral Large: —

Multilingual benchmarks
Benchmarkgpt-oss-20bPixtral Large
LMArena Non-English1268—
LMArena Chinese1314—
LMArena German1255—
LMArena Japanese1244—
LMArena Korean1236—
LMArena Russian1278—
LMArena Spanish1267—

Instruction Following Not comparable

gpt-oss-20b: 61.8 (#240), Pixtral Large: —

Instruction Following benchmarks
Benchmarkgpt-oss-20bPixtral Large
IFEval73.2%—
LMArena Instruction Following1236—

Long Context Not comparable

gpt-oss-20b: 37.9 (#209), Pixtral Large: —

Long Context benchmarks
Benchmarkgpt-oss-20bPixtral Large
LMArena Longer Query1250—

Writing & Preference gpt-oss-20b leads

gpt-oss-20b: 35.5 (#265), Pixtral Large: 32.9 (#278)

Writing & Preference benchmarks
Benchmarkgpt-oss-20bPixtral Large
EQ-Bench Creative Writing666988
LMArena Text1287—
LMArena Creative Writing1201—
WildBench73.7%—
LMArena Multi-Turn1268—

Frequently asked questions

Is gpt-oss-20b better than Pixtral Large?

gpt-oss-20b and Pixtral Large score almost the same on the Noometry Index (32.5 vs 32.2), so choose on price, context window or the category you care about most.

Which is cheaper, gpt-oss-20b or Pixtral Large?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Pixtral Large lists at $2 and $6.

Which has the bigger context window?

gpt-oss-20b does, with 131K tokens against 128K.

How many benchmarks do gpt-oss-20b and Pixtral Large share?

1 benchmark has published results for both models. gpt-oss-20b has 34 scored results on Noometry and Pixtral Large has 3.

Related comparisons

Go deeper