Model comparison

Claude Instant vs GPT-4

GPT-4 has enough public results to be ranked (#316); Claude Instant does not yet, so treat this comparison as directional.

Last verified . 6 shared benchmarks.

Claude Instant Anthropic

29.5

Unranked Sparse

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Claude Instant scores higher in 1 category and GPT-4 in 0 categories; one gap is clear of the uncertainty.
  • The biggest single-benchmark swing is DTBench: 45.8% for Claude Instant and 62.7% for GPT-4.

Side by side

Claude Instant and GPT-4 specifications
Claude InstantGPT-4
ProviderAnthropicOpenAI
Noometry Index29.529.1
Released2023-08-092023-03-14
WeightsProprietaryProprietary
Context window—8K
Max output—8K
Input $ / M tokens—$30
Output $ / M tokens—$60
Results tracked738

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude Instant: —, GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkClaude InstantGPT-4
HumanEval+50.6%79.3%
WeirdML—12.4%
BigCodeBench Instruct—46%
LMArena Coding—1254
BigCodeBench Complete—57.2%

Agentic & Tool Use Not comparable

Claude Instant: —, GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkClaude InstantGPT-4
METR Time Horizons—36.1%

Reasoning Claude Instant leads

Claude Instant: 19.6, GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkClaude InstantGPT-4
DTBench45.8%62.7%
Epoch Capabilities Index120.28125.89
Chess Puzzles—4%
LMArena Hard Prompts—1241
Mystery Game Puzzles—12%
LMCA—17.1%
BIG-Bench Hard—75.1%
ForecastBench—57.8
HellaSwag—95.3%
WinoGrande—87.5%

Math Not comparable

Claude Instant: —, GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkClaude InstantGPT-4
GSM8K86.7%92%
OTIS Mock AIME 2024-2025—1.1%
LMArena Math—1269
MATH Level 5—23%

Knowledge Not comparable

Claude Instant: —, GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkClaude InstantGPT-4
MMLU73.4%86.4%
TriviaQA78.9%84.8%
GPQA Diamond—35.7%
LMArena Expert—1211
ARC (AI2) Challenge86.3%—

Multilingual Not comparable

Claude Instant: —, GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkClaude InstantGPT-4
LMArena Non-English—1246
LMArena Chinese—1242
LMArena French—1283
LMArena German—1251
LMArena Japanese—1209
LMArena Korean—1184
LMArena Russian—1251
LMArena Spanish—1261

Instruction Following Not comparable

Claude Instant: —, GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkClaude InstantGPT-4
LMArena Instruction Following—1241

Long Context Not comparable

Claude Instant: —, GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkClaude InstantGPT-4
LMArena Longer Query—1244

Writing & Preference Not comparable

Claude Instant: —, GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkClaude InstantGPT-4
LMArena Text—1263
LMArena Creative Writing—1244
EQ-Bench Creative Writing—752
LMArena Multi-Turn—1257

Frequently asked questions

Is Claude Instant better than GPT-4?

GPT-4 has enough public results to be ranked (#316); Claude Instant does not yet, so treat this comparison as directional.

How many benchmarks do Claude Instant and GPT-4 share?

6 benchmarks have published results for both models. Claude Instant has 7 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper