Model comparison

GPT-5.3 Codex vs Qwen3.5 Max Preview

GPT-5.3 Codex and Qwen3.5 Max Preview score almost the same on the Noometry Index (45.8 vs 45.3), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

GPT-5.3 Codex OpenAI

45.8

Rank #69 Reported

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • The widest gap is in coding, where GPT-5.3 Codex leads 48.6 to 44.0.

Side by side

GPT-5.3 Codex and Qwen3.5 Max Preview specifications
GPT-5.3 CodexQwen3.5 Max Preview
ProviderOpenAIAlibaba (Qwen)
Noometry Index45.845.3
Released2026-02-05—
WeightsProprietaryProprietary
Context window400K—
Max output128K—
Input $ / M tokens$1.75—
Output $ / M tokens$14—
Results tracked817

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.3 Codex leads

GPT-5.3 Codex: 48.6 (#56), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkGPT-5.3 CodexQwen3.5 Max Preview
SWE-bench Verified74.8%—
LMArena WebDev1409—
WeirdML79.3%—
LMArena Coding—1487
ALE-Bench1,655—

Agentic & Tool Use Not comparable

GPT-5.3 Codex: 48.0 (#9), Qwen3.5 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.3 CodexQwen3.5 Max Preview
Terminal-Bench78.4%—
METR Time Horizons74.5%—
Vending-Bench 25,940—

Reasoning Not comparable

GPT-5.3 Codex: —, Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkGPT-5.3 CodexQwen3.5 Max Preview
LMArena Hard Prompts—1483
Epoch Capabilities Index156.77—

Math Not comparable

GPT-5.3 Codex: —, Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkGPT-5.3 CodexQwen3.5 Max Preview
LMArena Math—1474

Knowledge Not comparable

GPT-5.3 Codex: —, Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkGPT-5.3 CodexQwen3.5 Max Preview
LMArena Expert—1489

Multilingual Not comparable

GPT-5.3 Codex: —, Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkGPT-5.3 CodexQwen3.5 Max Preview
LMArena Non-English—1465
LMArena Chinese—1534
LMArena French—1484
LMArena German—1487
LMArena Japanese—1495
LMArena Korean—1438
LMArena Russian—1471
LMArena Spanish—1470

Instruction Following Not comparable

GPT-5.3 Codex: —, Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkGPT-5.3 CodexQwen3.5 Max Preview
LMArena Instruction Following—1467

Long Context Not comparable

GPT-5.3 Codex: —, Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkGPT-5.3 CodexQwen3.5 Max Preview
LMArena Longer Query—1476

Writing & Preference Not comparable

GPT-5.3 Codex: —, Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkGPT-5.3 CodexQwen3.5 Max Preview
LMArena Text—1470
LMArena Creative Writing—1464
LMArena Multi-Turn—1478

Frequently asked questions

Is GPT-5.3 Codex better than Qwen3.5 Max Preview?

GPT-5.3 Codex and Qwen3.5 Max Preview score almost the same on the Noometry Index (45.8 vs 45.3), so choose on price, context window or the category you care about most.

Is GPT-5.3 Codex or Qwen3.5 Max Preview better for coding?

GPT-5.3 Codex scores higher on coding benchmarks: 48.6 versus 44.0 in the Noometry coding category.

How many benchmarks do GPT-5.3 Codex and Qwen3.5 Max Preview share?

0 benchmarks have published results for both models. GPT-5.3 Codex has 8 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper