Model comparison

Chatgpt 4o Latest 20250326 vs Claude Sonnet 4.6

Claude Sonnet 4.6 is the stronger model overall, scoring 50.3 to 43.8 on the Noometry Index.

Last verified . 19 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Claude Sonnet 4.6 Anthropic

50.3

Rank #50 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 1 category and Claude Sonnet 4.6 in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 4.6 leads 52.9 to 38.6.

Side by side

Chatgpt 4o Latest 20250326 and Claude Sonnet 4.6 specifications
Chatgpt 4o Latest 20250326Claude Sonnet 4.6
ProviderOpenAIAnthropic
Noometry Index43.850.3
Released—2026-02-17
WeightsProprietaryProprietary
Context window—1M
Max output—128K
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked2157

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4.6 leads

Chatgpt 4o Latest 20250326: 41.6 (#122), Claude Sonnet 4.6: 46.3 (#67)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Sonnet 4.6
LMArena Coding14131504
SWE-bench Verified—75.2%
DeepSWE—29.9%
FrontierCode—24.3%
LMArena WebDev—1522
SciCode—46.8%
WeirdML—66.1%
ALE-Bench—1,327

Agentic & Tool Use Not comparable

Chatgpt 4o Latest 20250326: —, Claude Sonnet 4.6: 39.1 (#28)

Agentic & Tool Use benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Sonnet 4.6
Terminal-Bench—53.4%
APEX-Agents—43%
OSWorld 2.0—9.3%
DeepResearch Bench—54.9%
OSWorld—72.1%
ExploitBench—23.6%
GBAEval—48.8%
GDP.pdf—18%
LMArena Search—1221
Vending-Bench 2—7,204

Reasoning Claude Sonnet 4.6 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Claude Sonnet 4.6: 46.1 (#45)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Sonnet 4.6
LMArena Hard Prompts14241484
ARC-AGI-2—60.4%
Kagi LLM Benchmark75%—
NYT Connections (extended)—80.9%
ARC-AGI-1—86.5%
CritPt—3.1%
Chess Puzzles—13%
Thematic Generalization—76.3%
Mystery Game Puzzles—16%
DTBench—89.9%
LMCA—46.5%
Epoch Capabilities Index—152.24
ForecastBench—62

Math Claude Sonnet 4.6 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Claude Sonnet 4.6: 52.9 (#49)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Sonnet 4.6
LMArena Math14071462
OTIS Mock AIME 2024-2025—85.8%
ProofBench—45%
FrontierMath (Feb 2025 set)—32.4%
FrontierMath Tier 4 (v1)—8.3%

Knowledge Claude Sonnet 4.6 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Claude Sonnet 4.6: 51.7 (#65)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Sonnet 4.6
LMArena Expert14011500
GPQA Diamond—87.4%
SimpleQA Verified—35.5%
Confabulations16.6%—
Vectara Hallucination Rate—10.6%

Multimodal Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.6 (#58), Claude Sonnet 4.6: 38.0 (#68)

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Sonnet 4.6
LMArena Vision12431283
Blueprint-Bench 2—6.7%
LMArena Document—1482

Multilingual Claude Sonnet 4.6 leads

Chatgpt 4o Latest 20250326: 52.9 (#76), Claude Sonnet 4.6: 54.4 (#41)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Sonnet 4.6
LMArena Non-English14191440
LMArena Chinese14571491
LMArena French14461465
LMArena German14231428
LMArena Japanese14051420
LMArena Korean13961411
LMArena Russian14291440
LMArena Spanish14341464

Instruction Following Claude Sonnet 4.6 leads

Chatgpt 4o Latest 20250326: 74.0 (#107), Claude Sonnet 4.6: 77.4 (#25)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Sonnet 4.6
LMArena Instruction Following14031475

Long Context Claude Sonnet 4.6 leads

Chatgpt 4o Latest 20250326: 43.1 (#107), Claude Sonnet 4.6: 45.3 (#44)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Sonnet 4.6
LMArena Longer Query14131479

Writing & Preference Claude Sonnet 4.6 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Claude Sonnet 4.6: 70.2 (#22)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Claude Sonnet 4.6
LMArena Text14291458
LMArena Creative Writing14051435
EQ-Bench Creative Writing15011810
LMArena Multi-Turn14541464
EQ-Bench 4—1207

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Claude Sonnet 4.6?

Claude Sonnet 4.6 is the stronger model overall, scoring 50.3 to 43.8 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Claude Sonnet 4.6 better for coding?

Claude Sonnet 4.6 scores higher on coding benchmarks: 46.3 versus 41.6 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and Claude Sonnet 4.6 share?

19 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Claude Sonnet 4.6 has 57.

Related comparisons

Go deeper