Model comparison

GPT-3.5-turbo vs GPT-4

GPT-4 is the stronger model overall, scoring 29.1 to 23.2 on the Noometry Index. GPT-3.5-turbo costs 50× less per token, which makes it the better buy when GPT-4's lead doesn't matter for your workload.

Last verified . 37 shared benchmarks.

GPT-3.5-turbo OpenAI

23.2

Rank #350 Confirmed

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 37 benchmarks with published results for both. GPT-3.5-turbo scores higher in 0 categories and GPT-4 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where GPT-4 leads 34.9 to 25.3.
  • The biggest single-benchmark swing is DTBench: 48.5% for GPT-3.5-turbo and 62.7% for GPT-4.
  • GPT-3.5-turbo is cheaper at $0.50 / $1.50 per million input/output tokens, against $30 / $60 for GPT-4.
  • GPT-3.5-turbo accepts more context: 16K tokens versus 8K.

Side by side

GPT-3.5-turbo and GPT-4 specifications
GPT-3.5-turboGPT-4
ProviderOpenAIOpenAI
Noometry Index23.229.1
Released2023-03-012023-03-14
WeightsProprietaryProprietary
Context window16K8K
Max output4K8K
Input $ / M tokens$0.50$30
Output $ / M tokens$1.50$60
Results tracked4438

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4 leads

GPT-3.5-turbo: 23.9 (#331), GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkGPT-3.5-turboGPT-4
WeirdML3.5%12.4%
BigCodeBench Instruct39.1%46%
LMArena Coding11361254
BigCodeBench Complete50.6%57.2%
HumanEval+70.7%79.3%
MBPP+69.7%—

Agentic & Tool Use Not comparable

GPT-3.5-turbo: —, GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkGPT-3.5-turboGPT-4
METR Time Horizons21.5%36.1%

Reasoning GPT-4 leads

GPT-3.5-turbo: 13.8 (#332), GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkGPT-3.5-turboGPT-4
Chess Puzzles0%4%
LMArena Hard Prompts11081241
Mystery Game Puzzles3%12%
DTBench48.5%62.7%
LMCA9.7%17.1%
BIG-Bench Hard61.6%75.1%
Epoch Capabilities Index118.55125.89
ForecastBench50.457.8
WinoGrande81.6%87.5%
Adversarial NLI58.1%—
CommonsenseQA 2.057%—
HellaSwag—95.3%

Math GPT-4 leads

GPT-3.5-turbo: 6.3 (#327), GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkGPT-3.5-turboGPT-4
OTIS Mock AIME 2024-20252.2%1.1%
LMArena Math11421269
MATH Level 515.9%23%
GSM8K57.8%92%
FrontierMath (Tiers 1-3)0%—

Knowledge GPT-4 leads

GPT-3.5-turbo: 10.0 (#303), GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkGPT-3.5-turboGPT-4
GPQA Diamond28%35.7%
LMArena Expert10701211
MMLU71.4%86.4%
TriviaQA85.8%84.8%
ARC (AI2) Challenge87.4%—
BoolQ87%—
OpenBookQA86%—

Multilingual GPT-4 leads

GPT-3.5-turbo: 31.5 (#258), GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkGPT-3.5-turboGPT-4
LMArena Non-English11081246
LMArena Chinese10751242
LMArena French11181283
LMArena German10901251
LMArena Japanese10431209
LMArena Korean10191184
LMArena Russian11231251
LMArena Spanish11211261

Instruction Following GPT-4 leads

GPT-3.5-turbo: 57.9 (#262), GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkGPT-3.5-turboGPT-4
LMArena Instruction Following11191241

Long Context GPT-4 leads

GPT-3.5-turbo: 34.0 (#254), GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkGPT-3.5-turboGPT-4
LMArena Longer Query11211244

Writing & Preference GPT-4 leads

GPT-3.5-turbo: 25.3 (#305), GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkGPT-3.5-turboGPT-4
LMArena Text11251263
LMArena Creative Writing10921244
EQ-Bench Creative Writing451752
LMArena Multi-Turn11171257

Frequently asked questions

Is GPT-3.5-turbo better than GPT-4?

GPT-4 is the stronger model overall, scoring 29.1 to 23.2 on the Noometry Index. GPT-3.5-turbo costs 50× less per token, which makes it the better buy when GPT-4's lead doesn't matter for your workload.

Which is cheaper, GPT-3.5-turbo or GPT-4?

GPT-3.5-turbo is cheaper. It lists at $0.50 per million input tokens and $1.50 per million output tokens; GPT-4 lists at $30 and $60.

Is GPT-3.5-turbo or GPT-4 better for coding?

GPT-4 scores higher on coding benchmarks: 31.6 versus 23.9 in the Noometry coding category.

Which has the bigger context window?

GPT-3.5-turbo does, with 16K tokens against 8K.

How many benchmarks do GPT-3.5-turbo and GPT-4 share?

37 benchmarks have published results for both models. GPT-3.5-turbo has 44 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper