Model comparison

GPT-5.4 mini vs Qwen3.5 Max Preview

GPT-5.4 mini and Qwen3.5 Max Preview score almost the same on the Noometry Index (45.0 vs 45.3), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

GPT-5.4 mini OpenAI

45.0

Rank #76 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 17 benchmarks with published results for both. GPT-5.4 mini scores higher in 3 categories and Qwen3.5 Max Preview in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-5.4 mini leads 51.5 to 41.8.

Side by side

GPT-5.4 mini and Qwen3.5 Max Preview specifications
GPT-5.4 miniQwen3.5 Max Preview
ProviderOpenAIAlibaba (Qwen)
Noometry Index45.045.3
Released2026-03-17—
WeightsProprietaryProprietary
Context window400K—
Max output128K—
Input $ / M tokens$0.75—
Output $ / M tokens$4.50—
Results tracked4617

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 mini leads

GPT-5.4 mini: 45.2 (#72), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkGPT-5.4 miniQwen3.5 Max Preview
LMArena Coding14381487
FrontierCode27%—
LMArena WebDev1397—
SciCode49.9%—
WeirdML60.3%—
ALE-Bench1,189—

Agentic & Tool Use Not comparable

GPT-5.4 mini: 29.9 (#81), Qwen3.5 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.4 miniQwen3.5 Max Preview
DeepResearch Bench36.3%—

Reasoning Too close to call

GPT-5.4 mini: 30.4 (#85), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkGPT-5.4 miniQwen3.5 Max Preview
LMArena Hard Prompts14241483
ARC-AGI-218.9%—
Kagi LLM Benchmark37.9%—
NYT Connections (extended)61.8%—
ARC-AGI-163.7%—
CritPt10%—
Chess Puzzles24%—
Thematic Generalization61.7%—
Mystery Game Puzzles11%—
DTBench80%—
LMCA40.8%—
Epoch Capabilities Index148.84—
ForecastBench57—

Math GPT-5.4 mini leads

GPT-5.4 mini: 45.5 (#75), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkGPT-5.4 miniQwen3.5 Max Preview
LMArena Math14191474
FrontierMath (Tiers 1-3)51.2%—
FrontierMath Tier 49.8%—
OTIS Mock AIME 2024-202588.9%—
ProofBench21%—
FrontierMath (Feb 2025 set)28.3%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge GPT-5.4 mini leads

GPT-5.4 mini: 51.5 (#67), Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkGPT-5.4 miniQwen3.5 Max Preview
LMArena Expert14351489
GPQA Diamond86.9%—
SimpleQA Verified29.4%—
Vectara Hallucination Rate5.5%—

Multimodal Not comparable

GPT-5.4 mini: 39.7 (#56), Qwen3.5 Max Preview: —

Multimodal benchmarks
BenchmarkGPT-5.4 miniQwen3.5 Max Preview
LMArena Vision1245—

Multilingual Qwen3.5 Max Preview leads

GPT-5.4 mini: 51.9 (#96), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkGPT-5.4 miniQwen3.5 Max Preview
LMArena Non-English14051465
LMArena Chinese14461534
LMArena French14401484
LMArena German14091487
LMArena Japanese13741495
LMArena Korean13681438
LMArena Russian14171471
LMArena Spanish14051470

Instruction Following Qwen3.5 Max Preview leads

GPT-5.4 mini: 74.1 (#102), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkGPT-5.4 miniQwen3.5 Max Preview
LMArena Instruction Following14051467

Long Context Qwen3.5 Max Preview leads

GPT-5.4 mini: 43.0 (#112), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkGPT-5.4 miniQwen3.5 Max Preview
LMArena Longer Query14071476

Writing & Preference Qwen3.5 Max Preview leads

GPT-5.4 mini: 64.0 (#58), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkGPT-5.4 miniQwen3.5 Max Preview
LMArena Text14121470
LMArena Creative Writing13701464
LMArena Multi-Turn14291478
EQ-Bench Creative Writing1665—

Frequently asked questions

Is GPT-5.4 mini better than Qwen3.5 Max Preview?

GPT-5.4 mini and Qwen3.5 Max Preview score almost the same on the Noometry Index (45.0 vs 45.3), so choose on price, context window or the category you care about most.

Is GPT-5.4 mini or Qwen3.5 Max Preview better for coding?

GPT-5.4 mini scores higher on coding benchmarks: 45.2 versus 44.0 in the Noometry coding category.

How many benchmarks do GPT-5.4 mini and Qwen3.5 Max Preview share?

17 benchmarks have published results for both models. GPT-5.4 mini has 46 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper