Model comparison

GPT-5.2 Codex vs Qwen3.5 Plus

GPT-5.2 Codex and Qwen3.5 Plus score almost the same on the Noometry Index (42.6 vs 42.9), so choose on price, context window or the category you care about most.

Last verified . 1 shared benchmarks.

GPT-5.2 Codex OpenAI

42.6

Rank #111 Reported

Qwen3.5 Plus Alibaba (Qwen)

42.9

Rank #106 Confirmed

Summary

  • They share 1 benchmark with published results for both.
  • Qwen3.5 Plus is cheaper at $0.40 / $2.40 per million input/output tokens, against $1.75 / $14 for GPT-5.2 Codex.
  • Qwen3.5 Plus accepts more context: 1M tokens versus 400K.

Side by side

GPT-5.2 Codex and Qwen3.5 Plus specifications
GPT-5.2 CodexQwen3.5 Plus
ProviderOpenAIAlibaba (Qwen)
Noometry Index42.642.9
Released2025-12-182026-02-16
WeightsProprietaryProprietary
Context window400K1M
Max output128K66K
Input $ / M tokens$1.75$0.40
Output $ / M tokens$14$2.40
Results tracked515

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

GPT-5.2 Codex: 45.5 (#71), Qwen3.5 Plus: —

Coding benchmarks
BenchmarkGPT-5.2 CodexQwen3.5 Plus
ALE-Bench1,300621.92
SWE-bench Verified (bash only)72.8%—
LMArena WebDev1339—
SWE-bench Multilingual66.3%—

Agentic & Tool Use Not comparable

GPT-5.2 Codex: 41.0 (#22), Qwen3.5 Plus: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.2 CodexQwen3.5 Plus
Terminal-Bench66.5%—
Vending-Bench 2—0.54

Reasoning Not comparable

GPT-5.2 Codex: —, Qwen3.5 Plus: 32.8 (#74)

Reasoning benchmarks
BenchmarkGPT-5.2 CodexQwen3.5 Plus
Chess Puzzles—22%
Mystery Game Puzzles—17%
DTBench—80.5%
LMCA—36.4%
Epoch Capabilities Index—146.78

Math Not comparable

GPT-5.2 Codex: —, Qwen3.5 Plus: 49.6 (#61)

Math benchmarks
BenchmarkGPT-5.2 CodexQwen3.5 Plus
OTIS Mock AIME 2024-2025—86.7%
FrontierMath (Feb 2025 set)—21%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Not comparable

GPT-5.2 Codex: —, Qwen3.5 Plus: 46.0 (#83)

Knowledge benchmarks
BenchmarkGPT-5.2 CodexQwen3.5 Plus
GPQA Diamond—84.8%
SimpleQA Verified—25.4%
Vectara Hallucination Rate—10.7%

Long Context Not comparable

GPT-5.2 Codex: —, Qwen3.5 Plus: 43.0 (#113)

Long Context benchmarks
BenchmarkGPT-5.2 CodexQwen3.5 Plus
CL-bench—19.8%
CL-bench Life—12.4%

Frequently asked questions

Is GPT-5.2 Codex better than Qwen3.5 Plus?

GPT-5.2 Codex and Qwen3.5 Plus score almost the same on the Noometry Index (42.6 vs 42.9), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.2 Codex or Qwen3.5 Plus?

Qwen3.5 Plus is cheaper. It lists at $0.40 per million input tokens and $2.40 per million output tokens; GPT-5.2 Codex lists at $1.75 and $14.

Which has the bigger context window?

Qwen3.5 Plus does, with 1M tokens against 400K.

How many benchmarks do GPT-5.2 Codex and Qwen3.5 Plus share?

1 benchmark has published results for both models. GPT-5.2 Codex has 5 scored results on Noometry and Qwen3.5 Plus has 15.

Related comparisons

Go deeper