Model comparison

Claude Opus 4.5 vs GPT-6 Sol

GPT-6 Sol is the stronger model overall, scoring 61.8 to 50.5 on the Noometry Index.

Last verified . 37 shared benchmarks.

Claude Opus 4.5 Anthropic

50.5

Rank #47 Confirmed

GPT-6 Sol OpenAI

61.8

Rank #12 Confirmed

Summary

  • They share 37 benchmarks with published results for both. Claude Opus 4.5 scores higher in 4 categories and GPT-6 Sol in 6 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-6 Sol leads 87.2 to 38.6.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 4.9% for Claude Opus 4.5 and 90% for GPT-6 Sol.
  • GPT-6 Sol is cheaper at $2 / $10 per million input/output tokens, against $5 / $25 for Claude Opus 4.5.
  • GPT-6 Sol accepts more context: 1.05M tokens versus 200K.

Side by side

Claude Opus 4.5 and GPT-6 Sol specifications
Claude Opus 4.5GPT-6 Sol
ProviderAnthropicOpenAI
Noometry Index50.561.8
Released2025-11-012026-09-22
WeightsProprietaryProprietary
Context window200K1.05M
Max output64K128K
Input $ / M tokens$5$2
Output $ / M tokens$25$10
Results tracked6945

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-6 Sol leads

Claude Opus 4.5: 54.8 (#27), GPT-6 Sol: 60.1 (#11)

Coding benchmarks
BenchmarkClaude Opus 4.5GPT-6 Sol
LMArena WebDev14941688
LMArena Coding15041447
ALE-Bench1,0252,462
SWE-bench Verified76.7%—
DeepSWE—68.8%
FrontierCode—49.3%
SWE-bench Verified (bash only)76.8%—
SWE-bench Multilingual70.7%—
SciCode—57.6%
GSO26.5%—
WeirdML63.7%—
AlgoTune1.77—

Agentic & Tool Use Claude Opus 4.5 leads

Claude Opus 4.5: 47.3 (#12), GPT-6 Sol: 37.2 (#36)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.5GPT-6 Sol
Vending-Bench 24,96714,428
Terminal-Bench63.1%—
APEX-Agents—54.3%
Berkeley Function Calling Leaderboard77.5%—
GDPval45.5%—
Remote Labor Index3.8%—
τ²-bench Airline84%—
τ²-bench Banking24.7%—
τ²-bench Retail79.6%—
τ²-bench Telecom92.3%—
Cybench82%—
DeepResearch Bench54.8%—
OSWorld66.3%—
BALROG43.5%—
GDP.pdf—26.4%
LMArena Search1180—
METR Time Horizons75%—

Reasoning GPT-6 Sol leads

Claude Opus 4.5: 42.6 (#51), GPT-6 Sol: 74.0 (#9)

Reasoning benchmarks
BenchmarkClaude Opus 4.5GPT-6 Sol
ARC-AGI-237.6%89.6%
NYT Connections (extended)52.5%90.1%
ARC-AGI-180%95.5%
EBR-Bench14.3%53.3%
LMArena Hard Prompts14761418
Mystery Game Puzzles22%56%
DTBench89.9%97.3%
LMCA44.5%59.1%
Epoch Capabilities Index150.09162.72
SimpleBench62%—
Kagi LLM Benchmark80.2%—
CritPt—30.9%
Chess Puzzles12%—
EnigmaEval11.9%—
ForecastBench60.7—

Math GPT-6 Sol leads

Claude Opus 4.5: 38.6 (#132), GPT-6 Sol: 87.2 (#7)

Math benchmarks
BenchmarkClaude Opus 4.5GPT-6 Sol
FrontierMath (Tiers 1-3)34.4%89.8%
FrontierMath Tier 44.9%90%
OTIS Mock AIME 2024-202586.1%100%
ProofBench36%83%
LMArena Math14631402
FrontierMath (Feb 2025 set)20.7%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge GPT-6 Sol leads

Claude Opus 4.5: 56.5 (#44), GPT-6 Sol: 64.8 (#15)

Knowledge benchmarks
BenchmarkClaude Opus 4.5GPT-6 Sol
GPQA Diamond86%94.3%
SimpleQA Verified45.7%60.7%
Vectara Hallucination Rate10.9%6.5%
LMArena Expert14871439
Humanity's Last Exam25.2%—

Multimodal GPT-6 Sol leads

Claude Opus 4.5: 31.4 (#107), GPT-6 Sol: 47.6 (#10)

Multimodal benchmarks
BenchmarkClaude Opus 4.5GPT-6 Sol
Furniture Assembly28.3%58.3%
LMArena Vision—1245
GeoBench75%—
VPCT40%—
Blueprint-Bench 2—36.9%
LMArena Document1462—

Multilingual Claude Opus 4.5 leads

Claude Opus 4.5: 54.3 (#47), GPT-6 Sol: 50.5 (#118)

Multilingual benchmarks
BenchmarkClaude Opus 4.5GPT-6 Sol
LMArena Non-English14381385
LMArena Chinese14701405
LMArena French14711410
LMArena German14491390
LMArena Japanese14161385
LMArena Korean14241341
LMArena Russian14471401
LMArena Spanish14581384

Instruction Following Claude Opus 4.5 leads

Claude Opus 4.5: 77.5 (#19), GPT-6 Sol: 74.5 (#94)

Instruction Following benchmarks
BenchmarkClaude Opus 4.5GPT-6 Sol
LMArena Instruction Following14781412

Long Context Claude Opus 4.5 leads

Claude Opus 4.5: 46.5 (#22), GPT-6 Sol: 43.1 (#108)

Long Context benchmarks
BenchmarkClaude Opus 4.5GPT-6 Sol
LMArena Longer Query14801411
CL-bench21.1%—

Writing & Preference GPT-6 Sol leads

Claude Opus 4.5: 68.1 (#28), GPT-6 Sol: 71.9 (#18)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.5GPT-6 Sol
LMArena Text14511395
LMArena Creative Writing14451378
EQ-Bench Creative Writing16872125
LMArena Multi-Turn14661412

Frequently asked questions

Is Claude Opus 4.5 better than GPT-6 Sol?

GPT-6 Sol is the stronger model overall, scoring 61.8 to 50.5 on the Noometry Index.

Which is cheaper, Claude Opus 4.5 or GPT-6 Sol?

GPT-6 Sol is cheaper. It lists at $2 per million input tokens and $10 per million output tokens; Claude Opus 4.5 lists at $5 and $25.

Is Claude Opus 4.5 or GPT-6 Sol better for coding?

GPT-6 Sol scores higher on coding benchmarks: 60.1 versus 54.8 in the Noometry coding category.

Which has the bigger context window?

GPT-6 Sol does, with 1.05M tokens against 200K.

How many benchmarks do Claude Opus 4.5 and GPT-6 Sol share?

37 benchmarks have published results for both models. Claude Opus 4.5 has 69 scored results on Noometry and GPT-6 Sol has 45.

Related comparisons

Go deeper