Model comparison

Claude 3.7 Sonnet vs GPT-5.1-Codex

Claude 3.7 Sonnet and GPT-5.1-Codex score almost the same on the Noometry Index (39.5 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

GPT-5.1-Codex OpenAI

38.6

Rank #186 Reported

Summary

  • They share 2 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 1 category and GPT-5.1-Codex in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude 3.7 Sonnet leads 37.5 to 30.3.
  • The biggest single-benchmark swing is SWE-bench Verified (bash only): 52.8% for Claude 3.7 Sonnet and 66% for GPT-5.1-Codex.

Side by side

Claude 3.7 Sonnet and GPT-5.1-Codex specifications
Claude 3.7 SonnetGPT-5.1-Codex
ProviderAnthropicOpenAI
Noometry Index39.538.6
Released2025-02-242025-11-12
WeightsProprietaryProprietary
Context window—400K
Max output—128K
Input $ / M tokens—$1.25
Output $ / M tokens—$10
Results tracked586

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.1-Codex leads

Claude 3.7 Sonnet: 40.6 (#136), GPT-5.1-Codex: 41.9 (#116)

Coding benchmarks
BenchmarkClaude 3.7 SonnetGPT-5.1-Codex
SWE-bench Verified (bash only)52.8%66%
SWE-bench Verified61%—
Aider Polyglot64.9%—
LMArena WebDev—1337
GSO3.8%—
LiveBench Coding74.5%—
LMArena Coding1361—
CadEval54%—
ALE-Bench—1,245

Agentic & Tool Use GPT-5.1-Codex leads

Claude 3.7 Sonnet: 34.1 (#50), GPT-5.1-Codex: 38.0 (#33)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetGPT-5.1-Codex
METR Time Horizons60%70.8%
Terminal-Bench—60.4%
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—

Reasoning Not comparable

Claude 3.7 Sonnet: 18.6 (#277), GPT-5.1-Codex: —

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetGPT-5.1-Codex
ARC-AGI-20.9%—
SimpleBench46.4%—
ARC-AGI-128.6%—
EnigmaEval4.2%—
LiveBench Reasoning87.8%—
LMArena Hard Prompts1333—
LiveBench Data Analysis74%—
Epoch Capabilities Index141.16—
ForecastBench61.8—
LiveBench76.1%—

Math Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 37.5 (#153), GPT-5.1-Codex: 30.3 (#235)

Math benchmarks
BenchmarkClaude 3.7 SonnetGPT-5.1-Codex
OTIS Mock AIME 2024-202557.8%—
ProofBench—9%
Omni-MATH33%—
LiveBench Math79%—
LMArena Math1337—
MATH Level 591.2%—
FrontierMath (Feb 2025 set)4.1%—

Knowledge Not comparable

Claude 3.7 Sonnet: 39.8 (#130), GPT-5.1-Codex: —

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetGPT-5.1-Codex
GPQA Diamond79.7%—
Humanity's Last Exam8%—
MMLU-Pro78.4%—
Confabulations14.7%—
GPQA (HELM)60.8%—
LMArena Expert1321—

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), GPT-5.1-Codex: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetGPT-5.1-Codex
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Not comparable

Claude 3.7 Sonnet: 44.1 (#179), GPT-5.1-Codex: —

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetGPT-5.1-Codex
LMArena Non-English1296—
LMArena Chinese1299—
LMArena French1303—
LMArena German1301—
LMArena Japanese1267—
LMArena Korean1249—
LMArena Russian1311—
LMArena Spanish1298—

Instruction Following Not comparable

Claude 3.7 Sonnet: 72.9 (#125), GPT-5.1-Codex: —

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetGPT-5.1-Codex
LiveBench Instruction Following81.3%—
IFEval83.4%—
LMArena Instruction Following1352—

Long Context Not comparable

Claude 3.7 Sonnet: 50.3 (#10), GPT-5.1-Codex: —

Long Context benchmarks
BenchmarkClaude 3.7 SonnetGPT-5.1-Codex
Fiction.LiveBench83.3%—
LMArena Longer Query1373—

Writing & Preference Not comparable

Claude 3.7 Sonnet: 54.4 (#150), GPT-5.1-Codex: —

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetGPT-5.1-Codex
LMArena Text1314—
LMArena Creative Writing1332—
Short-Story Creative Writing81.1%—
EQ-Bench Creative Writing1412—
WildBench81.4%—
LMArena Multi-Turn1339—
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than GPT-5.1-Codex?

Claude 3.7 Sonnet and GPT-5.1-Codex score almost the same on the Noometry Index (39.5 vs 38.6), so choose on price, context window or the category you care about most.

Is Claude 3.7 Sonnet or GPT-5.1-Codex better for coding?

GPT-5.1-Codex scores higher on coding benchmarks: 41.9 versus 40.6 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and GPT-5.1-Codex share?

2 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and GPT-5.1-Codex has 6.

Related comparisons

Go deeper