Model comparison

GPT-5.4 nano vs o1-mini

GPT-5.4 nano is the stronger model overall, scoring 41.9 to 34.0 on the Noometry Index.

Last verified . 24 shared benchmarks.

GPT-5.4 nano OpenAI

41.9

Rank #125 Confirmed

o1-mini OpenAI

34.0

Rank #235 Confirmed

Summary

  • They share 24 benchmarks with published results for both. GPT-5.4 nano scores higher in 8 categories and o1-mini in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-5.4 nano leads 23.7 to 8.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 87.8% for GPT-5.4 nano and 46.9% for o1-mini.

Side by side

GPT-5.4 nano and o1-mini specifications
GPT-5.4 nanoo1-mini
ProviderOpenAIOpenAI
Noometry Index41.934.0
Released2026-03-172024-09-12
WeightsProprietaryProprietary
Context window400K—
Max output128K—
Input $ / M tokens$0.20—
Output $ / M tokens$1.25—
Results tracked4039

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 nano leads

GPT-5.4 nano: 43.6 (#84), o1-mini: 35.5 (#224)

Coding benchmarks
BenchmarkGPT-5.4 nanoo1-mini
WeirdML49.2%36.3%
LMArena Coding14051362
Aider Polyglot—32.9%
SciCode46.9%—
LiveBench Coding—48%
ALE-Bench1,005—
HumanEval+—89%
MBPP+—78.8%

Agentic & Tool Use Not comparable

GPT-5.4 nano: —, o1-mini: 24.6 (#118)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.4 nanoo1-mini
Cybench—10%

Reasoning GPT-5.4 nano leads

GPT-5.4 nano: 23.7 (#173), o1-mini: 8.8 (#346)

Reasoning benchmarks
BenchmarkGPT-5.4 nanoo1-mini
ARC-AGI-25.7%0.8%
ARC-AGI-151.5%14%
LMArena Hard Prompts13811333
Epoch Capabilities Index145.81135.82
SimpleBench—18.1%
Kagi LLM Benchmark39.7%—
CritPt9.3%—
Chess Puzzles30%—
LiveBench Reasoning—72.3%
Mystery Game Puzzles9%—
DTBench80.3%—
LiveBench Data Analysis—57.9%
LMCA36.9%—
ForecastBench57.3—
LiveBench—57.8%

Math GPT-5.4 nano leads

GPT-5.4 nano: 40.9 (#88), o1-mini: 35.4 (#186)

Math benchmarks
BenchmarkGPT-5.4 nanoo1-mini
OTIS Mock AIME 2024-202587.8%46.9%
LMArena Math14061358
FrontierMath (Feb 2025 set)25.9%1.7%
FrontierMath (Tiers 1-3)44.9%—
FrontierMath Tier 412.2%—
ProofBench5%—
LiveBench Math—62%
MATH Level 5—89.2%
FrontierMath Tier 4 (v1)6.3%—

Knowledge GPT-5.4 nano leads

GPT-5.4 nano: 41.9 (#103), o1-mini: 34.9 (#192)

Knowledge benchmarks
BenchmarkGPT-5.4 nanoo1-mini
GPQA Diamond78.5%62.4%
LMArena Expert13961316
SimpleQA Verified11.7%—
Confabulations—18.6%
Vectara Hallucination Rate3.1%—

Multimodal Not comparable

GPT-5.4 nano: 36.7 (#78), o1-mini: —

Multimodal benchmarks
BenchmarkGPT-5.4 nanoo1-mini
LMArena Vision1196—

Multilingual GPT-5.4 nano leads

GPT-5.4 nano: 48.6 (#140), o1-mini: 43.6 (#182)

Multilingual benchmarks
BenchmarkGPT-5.4 nanoo1-mini
LMArena Non-English13591289
LMArena Chinese13921314
LMArena French13961293
LMArena German13671278
LMArena Japanese13431245
LMArena Korean13201223
LMArena Russian13631283
LMArena Spanish13711303

Instruction Following GPT-5.4 nano leads

GPT-5.4 nano: 71.9 (#144), o1-mini: 66.7 (#206)

Instruction Following benchmarks
BenchmarkGPT-5.4 nanoo1-mini
LMArena Instruction Following13621304
LiveBench Instruction Following—65.4%

Long Context GPT-5.4 nano leads

GPT-5.4 nano: 41.6 (#137), o1-mini: 40.1 (#161)

Long Context benchmarks
BenchmarkGPT-5.4 nanoo1-mini
LMArena Longer Query13661320

Writing & Preference GPT-5.4 nano leads

GPT-5.4 nano: 55.7 (#142), o1-mini: 48.4 (#202)

Writing & Preference benchmarks
BenchmarkGPT-5.4 nanoo1-mini
LMArena Text13721317
LMArena Creative Writing13141244
LMArena Multi-Turn13821314
Short-Story Creative Writing—64.9%
LiveBench Language—40.9%

Frequently asked questions

Is GPT-5.4 nano better than o1-mini?

GPT-5.4 nano is the stronger model overall, scoring 41.9 to 34.0 on the Noometry Index.

Is GPT-5.4 nano or o1-mini better for coding?

GPT-5.4 nano scores higher on coding benchmarks: 43.6 versus 35.5 in the Noometry coding category.

How many benchmarks do GPT-5.4 nano and o1-mini share?

24 benchmarks have published results for both models. GPT-5.4 nano has 40 scored results on Noometry and o1-mini has 39.

Related comparisons

Go deeper