Model comparison

GPT-4.5 vs Qwen3 Max

Qwen3 Max is the stronger model overall, scoring 43.7 to 37.2 on the Noometry Index.

Last verified . 21 shared benchmarks.

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 21 benchmarks with published results for both. GPT-4.5 scores higher in 0 categories and Qwen3 Max in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 Max leads 48.1 to 32.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 37.8% for GPT-4.5 and 73.3% for Qwen3 Max.

Side by side

GPT-4.5 and Qwen3 Max specifications
GPT-4.5Qwen3 Max
ProviderOpenAIAlibaba (Qwen)
Noometry Index37.243.7
Released2025-02-272025-09-23
WeightsProprietaryProprietary
Context window—262K
Max output—66K
Input $ / M tokens—$1.20
Output $ / M tokens—$6
Results tracked4233

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-4.5: 42.2 (#109), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkGPT-4.5Qwen3 Max
LMArena Coding13961456
Aider Polyglot44.9%—
WeirdML39.4%—
LiveBench Coding75.2%—
ALE-Bench—370.45

Agentic & Tool Use Not comparable

GPT-4.5: 27.9 (#97), Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4.5Qwen3 Max
Cybench17.5%—
Vending-Bench 2—71.56

Reasoning Qwen3 Max leads

GPT-4.5: 13.9 (#330), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkGPT-4.5Qwen3 Max
LMArena Hard Prompts14031448
Epoch Capabilities Index136.74142.38
ARC-AGI-20.8%—
SimpleBench34.5%—
Kagi LLM Benchmark—72.5%
NYT Connections (extended)—30.1%
ARC-AGI-110.3%—
Chess Puzzles—4%
EnigmaEval3.2%—
LiveBench Reasoning71.1%—
Mystery Game Puzzles—5%
DTBench—82.1%
LiveBench Data Analysis64.3%—
LMCA—28.3%
ForecastBench61.7—
LiveBench69%—

Math Qwen3 Max leads

GPT-4.5: 32.6 (#211), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkGPT-4.5Qwen3 Max
OTIS Mock AIME 2024-202537.8%73.3%
LMArena Math14121446
MATH Level 578.6%97.1%
FrontierMath (Tiers 1-3)—18.9%
LiveBench Math69.3%—

Knowledge Qwen3 Max leads

GPT-4.5: 32.5 (#211), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkGPT-4.5Qwen3 Max
GPQA Diamond68.7%72.6%
LMArena Expert13941455
Humanity's Last Exam5.4%—
SimpleQA Verified—48.7%
Confabulations13.6%—

Multimodal Not comparable

GPT-4.5: 37.6 (#71), Qwen3 Max: —

Multimodal benchmarks
BenchmarkGPT-4.5Qwen3 Max
LMArena Vision1195—
VPCT45%—

Multilingual Qwen3 Max leads

GPT-4.5: 52.5 (#83), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkGPT-4.5Qwen3 Max
LMArena Non-English14131429
LMArena Chinese14211478
LMArena French14181449
LMArena German14571463
LMArena Japanese14161397
LMArena Korean13921399
LMArena Russian14191428
LMArena Spanish—1462

Instruction Following Qwen3 Max leads

GPT-4.5: 72.6 (#134), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkGPT-4.5Qwen3 Max
LMArena Instruction Following14041419
LiveBench Instruction Following72.3%—

Long Context Qwen3 Max leads

GPT-4.5: 40.4 (#155), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkGPT-4.5Qwen3 Max
Fiction.LiveBench63.9%66.7%
LMArena Longer Query14061438
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

GPT-4.5: 56.9 (#134), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkGPT-4.5Qwen3 Max
LMArena Text14171439
LMArena Creative Writing13941402
LMArena Multi-Turn14441446
Short-Story Creative Writing75.6%—
EQ-Bench Creative Writing1258—
LiveBench Language61.5%—

Frequently asked questions

Is GPT-4.5 better than Qwen3 Max?

Qwen3 Max is the stronger model overall, scoring 43.7 to 37.2 on the Noometry Index.

Is GPT-4.5 or Qwen3 Max better for coding?

They score almost the same on coding (42.2 vs 43.0); test both on your own repository before choosing.

How many benchmarks do GPT-4.5 and Qwen3 Max share?

21 benchmarks have published results for both models. GPT-4.5 has 42 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper