Model comparison

Claude 3.5 Sonnet vs GPT-4.1

GPT-4.1 is the stronger model overall, scoring 35.9 to 34.6 on the Noometry Index.

Last verified . 39 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

Summary

  • They share 39 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 2 categories and GPT-4.1 in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in multimodal, where GPT-4.1 leads 38.2 to 26.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.5% for Claude 3.5 Sonnet and 38.3% for GPT-4.1.

Side by side

Claude 3.5 Sonnet and GPT-4.1 specifications
Claude 3.5 SonnetGPT-4.1
ProviderAnthropicOpenAI
Noometry Index34.635.9
Released2024-06-202025-04-14
WeightsProprietaryProprietary
Context window—1.05M
Max output—33K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked6052

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), GPT-4.1: 34.4 (#238)

Coding benchmarks
BenchmarkClaude 3.5 SonnetGPT-4.1
Aider Polyglot51.6%52.4%
WeirdML40%39%
LMArena Coding13421391
CadEval48%42%
SWE-bench Verified—48.5%
SWE-bench Verified (bash only)—39.6%
GSO4.6%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
ALE-Bench—558.1
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use GPT-4.1 leads

Claude 3.5 Sonnet: 32.3 (#67), GPT-4.1: 34.7 (#43)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetGPT-4.1
Berkeley Function Calling Leaderboard—54%
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 23.1 (#183), GPT-4.1: 11.7 (#339)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetGPT-4.1
SimpleBench41.4%27%
EnigmaEval0.9%2.2%
LMArena Hard Prompts13051384
DTBench67.8%68.3%
Epoch Capabilities Index133.55136.78
ForecastBench60.761.5
ARC-AGI-2—0.4%
Kagi LLM Benchmark—52.3%
ARC-AGI-1—5.5%
Chess Puzzles—6%
LiveBench Reasoning56.7%—
LiveBench Data Analysis55%—
LMCA—25.6%
LiveBench59%—

Math GPT-4.1 leads

Claude 3.5 Sonnet: 19.2 (#288), GPT-4.1: 22.3 (#280)

Math benchmarks
BenchmarkClaude 3.5 SonnetGPT-4.1
OTIS Mock AIME 2024-20258.5%38.3%
Omni-MATH27.6%47.1%
LMArena Math13071370
MATH Level 556.9%83%
FrontierMath (Feb 2025 set)2.1%5.5%
FrontierMath Tier 4 (v1)0%0%
FrontierMath (Tiers 1-3)—6%
LiveBench Math52.3%—

Knowledge GPT-4.1 leads

Claude 3.5 Sonnet: 28.6 (#245), GPT-4.1: 37.1 (#160)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetGPT-4.1
GPQA Diamond55.3%66.9%
Humanity's Last Exam4.1%5.4%
MMLU-Pro77.7%81.1%
GPQA (HELM)56.5%65.9%
LMArena Expert12651364
SimpleQA Verified—31.1%
Confabulations19.9%—
Vectara Hallucination Rate—5.6%
MMLU87.3%—

Multimodal GPT-4.1 leads

Claude 3.5 Sonnet: 26.5 (#120), GPT-4.1: 38.2 (#67)

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetGPT-4.1
LMArena Vision11251211
GeoBench62%72%
Video-MME60%—
VPCT33%—

Multilingual GPT-4.1 leads

Claude 3.5 Sonnet: 43.2 (#185), GPT-4.1: 49.4 (#133)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetGPT-4.1
LMArena Non-English12831370
LMArena Chinese12721382
LMArena French13051382
LMArena German12971381
LMArena Japanese12341319
LMArena Korean12001339
LMArena Russian13061377
LMArena Spanish12901376

Instruction Following GPT-4.1 leads

Claude 3.5 Sonnet: 68.8 (#182), GPT-4.1: 71.3 (#153)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetGPT-4.1
IFEval85.5%83.8%
LMArena Instruction Following12971367
LiveBench Instruction Following69.3%—

Long Context Too close to call

Claude 3.5 Sonnet: 39.9 (#167), GPT-4.1: 40.0 (#163)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetGPT-4.1
LMArena Longer Query13111385
Fiction.LiveBench—63.9%

Writing & Preference GPT-4.1 leads

Claude 3.5 Sonnet: 52.9 (#164), GPT-4.1: 57.6 (#125)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetGPT-4.1
LMArena Text12981383
LMArena Creative Writing12921363
EQ-Bench Creative Writing14511420
WildBench79.2%85.4%
LMArena Multi-Turn13261398
Short-Story Creative Writing80.3%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than GPT-4.1?

GPT-4.1 is the stronger model overall, scoring 35.9 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or GPT-4.1 better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 34.4 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and GPT-4.1 share?

39 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and GPT-4.1 has 52.

Related comparisons

Go deeper