Model comparison

GPT-5.1 vs GPT-5 Mini

GPT-5.1 is the stronger model overall, scoring 49.0 to 41.8 on the Noometry Index. GPT-5 Mini costs 5.0× less per token, which makes it the better buy when GPT-5.1's lead doesn't matter for your workload.

Last verified . 48 shared benchmarks.

GPT-5.1 OpenAI

49.0

Rank #53 Confirmed

GPT-5 Mini OpenAI

41.8

Rank #128 Confirmed

Summary

  • They share 48 benchmarks with published results for both. GPT-5.1 scores higher in 10 categories and GPT-5 Mini in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-5.1 leads 39.8 to 23.9.
  • The biggest single-benchmark swing is GPQA (HELM): 44.2% for GPT-5.1 and 75.6% for GPT-5 Mini.
  • GPT-5 Mini is cheaper at $0.25 / $2 per million input/output tokens, against $1.25 / $10 for GPT-5.1.

Side by side

GPT-5.1 and GPT-5 Mini specifications
GPT-5.1GPT-5 Mini
ProviderOpenAIOpenAI
Noometry Index49.041.8
Released2025-11-132025-08-07
WeightsProprietaryProprietary
Context window400K400K
Max output128K128K
Input $ / M tokens$1.25$0.25
Output $ / M tokens$10$2
Results tracked6360

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.1 leads

GPT-5.1: 46.4 (#66), GPT-5 Mini: 40.1 (#146)

Coding benchmarks
BenchmarkGPT-5.1GPT-5 Mini
SWE-bench Verified68%64.7%
SWE-bench Verified (bash only)66%59.8%
SciCode43.3%39.2%
WeirdML60.8%52.7%
LMArena Coding14541406
ALE-Bench1,192799.77
LMArena WebDev1395—
SWE-bench Multilingual—39.7%
GSO13.7%—
LiveBench Coding72.5%—
AlgoTune—1.38

Agentic & Tool Use GPT-5.1 leads

GPT-5.1: 32.7 (#60), GPT-5 Mini: 31.1 (#70)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.1GPT-5 Mini
Terminal-Bench47.6%34.8%
Vending-Bench 21,473-31.18
Berkeley Function Calling Leaderboard—55.5%
DeepResearch Bench42.8%—
LMArena Search1199—

Reasoning GPT-5.1 leads

GPT-5.1: 39.8 (#58), GPT-5 Mini: 23.9 (#168)

Reasoning benchmarks
BenchmarkGPT-5.1GPT-5 Mini
ARC-AGI-217.6%4.4%
ARC-AGI-172.8%54.3%
CritPt4.9%0%
Chess Puzzles32%30%
EnigmaEval11.2%8.2%
LMArena Hard Prompts14571380
Mystery Game Puzzles19%10%
DTBench90.1%80.5%
LMCA43.9%34.2%
Epoch Capabilities Index149.64145.52
ForecastBench58.161
SimpleBench53.2%—
Kagi LLM Benchmark—70.3%
LiveBench Reasoning95.8%—
LiveBench Data Analysis72.1%—
LiveBench78.8%—

Math GPT-5.1 leads

GPT-5.1: 52.2 (#51), GPT-5 Mini: 46.7 (#69)

Math benchmarks
BenchmarkGPT-5.1GPT-5 Mini
OTIS Mock AIME 2024-202588.6%86.7%
Omni-MATH46.4%72.2%
LMArena Math14471378
FrontierMath (Feb 2025 set)31%27.2%
FrontierMath Tier 4 (v1)12.5%6.3%
FrontierMath (Tiers 1-3)—46.7%
FrontierMath Tier 4—12.2%
ProofBench—9%
LiveBench Math94.5%—
MATH Level 5—97.8%

Knowledge GPT-5.1 leads

GPT-5.1: 50.6 (#71), GPT-5 Mini: 45.6 (#86)

Knowledge benchmarks
BenchmarkGPT-5.1GPT-5 Mini
GPQA Diamond87.6%75%
Humanity's Last Exam23.7%19.4%
SimpleQA Verified48%21.6%
MMLU-Pro57.9%83.5%
Vectara Hallucination Rate10.9%12.9%
GPQA (HELM)44.2%75.6%
LMArena Expert14701379
Confabulations—13.3%

Multimodal GPT-5.1 leads

GPT-5.1: 44.8 (#19), GPT-5 Mini: 35.6 (#85)

Multimodal benchmarks
BenchmarkGPT-5.1GPT-5 Mini
LMArena Vision12501202
VPCT58.7%40.2%
LMArena Document1403—

Multilingual GPT-5.1 leads

GPT-5.1: 53.8 (#56), GPT-5 Mini: 48.9 (#137)

Multilingual benchmarks
BenchmarkGPT-5.1GPT-5 Mini
LMArena Non-English14311363
LMArena Chinese14951385
LMArena French14501386
LMArena German14381366
LMArena Japanese14531341
LMArena Korean14011308
LMArena Russian14351362
LMArena Spanish14331355

Instruction Following GPT-5.1 leads

GPT-5.1: 83.9 (#1), GPT-5 Mini: 76.2 (#46)

Instruction Following benchmarks
BenchmarkGPT-5.1GPT-5 Mini
IFEval93.5%92.7%
LMArena Instruction Following14431357
LiveBench Instruction Following93.3%—

Long Context GPT-5.1 leads

GPT-5.1: 47.6 (#14), GPT-5 Mini: 41.9 (#132)

Long Context benchmarks
BenchmarkGPT-5.1GPT-5 Mini
LMArena Longer Query14471355
Fiction.LiveBench—69.4%
CL-bench23.7%—
CL-bench Life17.3%—

Writing & Preference GPT-5.1 leads

GPT-5.1: 64.5 (#55), GPT-5 Mini: 55.2 (#148)

Writing & Preference benchmarks
BenchmarkGPT-5.1GPT-5 Mini
LMArena Text14431373
LMArena Creative Writing14271325
WildBench86.3%85.5%
LMArena Multi-Turn14501363
Short-Story Creative Writing—83.1%
EQ-Bench Creative Writing—1313
LiveBench Language80.2%—

Frequently asked questions

Is GPT-5.1 better than GPT-5 Mini?

GPT-5.1 is the stronger model overall, scoring 49.0 to 41.8 on the Noometry Index. GPT-5 Mini costs 5.0× less per token, which makes it the better buy when GPT-5.1's lead doesn't matter for your workload.

Which is cheaper, GPT-5.1 or GPT-5 Mini?

GPT-5 Mini is cheaper. It lists at $0.25 per million input tokens and $2 per million output tokens; GPT-5.1 lists at $1.25 and $10.

Is GPT-5.1 or GPT-5 Mini better for coding?

GPT-5.1 scores higher on coding benchmarks: 46.4 versus 40.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 400K tokens.

How many benchmarks do GPT-5.1 and GPT-5 Mini share?

48 benchmarks have published results for both models. GPT-5.1 has 63 scored results on Noometry and GPT-5 Mini has 60.

Related comparisons

Go deeper