Model comparison

GPT-4 Turbo vs o1

o1 is the stronger model overall, scoring 40.9 to 30.5 on the Noometry Index. GPT-4 Turbo costs 1.8× less per token, which makes it the better buy when o1's lead doesn't matter for your workload.

Last verified . 32 shared benchmarks.

GPT-4 Turbo OpenAI

30.5

Rank #292 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 32 benchmarks with published results for both. GPT-4 Turbo scores higher in 0 categories and o1 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where o1 leads 36.1 to 9.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 6.7% for GPT-4 Turbo and 73.3% for o1.
  • GPT-4 Turbo is cheaper at $10 / $30 per million input/output tokens, against $15 / $60 for o1.
  • o1 accepts more context: 200K tokens versus 128K.

Side by side

GPT-4 Turbo and o1 specifications
GPT-4 Turboo1
ProviderOpenAIOpenAI
Noometry Index30.540.9
Released2023-11-062024-09-12
WeightsProprietaryProprietary
Context window128K200K
Max output4K100K
Input $ / M tokens$10$15
Output $ / M tokens$30$60
Results tracked3652

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

GPT-4 Turbo: 33.8 (#249), o1: 46.1 (#70)

Coding benchmarks
BenchmarkGPT-4 Turboo1
WeirdML18%47.6%
LMArena Coding12681367
HumanEval+86.6%89%
MBPP+73.3%80.2%
Aider Polyglot—61.7%
BigCodeBench Instruct48.2%—
LiveBench Coding—69.7%
BigCodeBench Complete58.2%—
CadEval—56%

Agentic & Tool Use Not comparable

GPT-4 Turbo: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkGPT-4 Turboo1
METR Time Horizons36.7%51.1%
Cybench—10%

Reasoning o1 leads

GPT-4 Turbo: 15.3 (#317), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkGPT-4 Turboo1
SimpleBench25.1%41.7%
Chess Puzzles6%15%
LMArena Hard Prompts12511371
DTBench61.6%74.7%
LMCA9.8%22.3%
Epoch Capabilities Index127.25141.91
ARC-AGI-1—30.7%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
LiveBench Data Analysis—65.5%
ForecastBench59.4—
LiveBench—75.7%

Math o1 leads

GPT-4 Turbo: 9.0 (#322), o1: 36.1 (#175)

Math benchmarks
BenchmarkGPT-4 Turboo1
FrontierMath (Tiers 1-3)0.7%14.7%
OTIS Mock AIME 2024-20256.7%73.3%
LMArena Math12721388
MATH Level 546.7%94.7%
LiveBench Math—80.3%
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

GPT-4 Turbo: 24.3 (#268), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkGPT-4 Turboo1
GPQA Diamond46.6%76.8%
Confabulations28.4%11.7%
LMArena Expert12231361
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
MMLU81.3%—

Multimodal o1 leads

GPT-4 Turbo: 30.6 (#110), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkGPT-4 Turboo1
LMArena Vision10901168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual o1 leads

GPT-4 Turbo: 40.5 (#216), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkGPT-4 Turboo1
LMArena Non-English12451358
LMArena Chinese12421394
LMArena French12761344
LMArena German12591337
LMArena Japanese11941346
LMArena Korean11871396
LMArena Russian12591356
LMArena Spanish12601345

Instruction Following o1 leads

GPT-4 Turbo: 65.8 (#216), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkGPT-4 Turboo1
LMArena Instruction Following12491367
LiveBench Instruction Following—81.5%

Long Context o1 leads

GPT-4 Turbo: 38.0 (#206), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkGPT-4 Turboo1
LMArena Longer Query12541378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

GPT-4 Turbo: 47.7 (#206), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkGPT-4 Turboo1
LMArena Text12721366
LMArena Creative Writing12691348
LMArena Multi-Turn12671369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is GPT-4 Turbo better than o1?

o1 is the stronger model overall, scoring 40.9 to 30.5 on the Noometry Index. GPT-4 Turbo costs 1.8× less per token, which makes it the better buy when o1's lead doesn't matter for your workload.

Which is cheaper, GPT-4 Turbo or o1?

GPT-4 Turbo is cheaper. It lists at $10 per million input tokens and $30 per million output tokens; o1 lists at $15 and $60.

Is GPT-4 Turbo or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 33.8 in the Noometry coding category.

Which has the bigger context window?

o1 does, with 200K tokens against 128K.

How many benchmarks do GPT-4 Turbo and o1 share?

32 benchmarks have published results for both models. GPT-4 Turbo has 36 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper