Model comparison

Qwen Max vs Qwen3 Max

Qwen3 Max is the stronger model overall, scoring 43.7 to 34.7 on the Noometry Index.

Last verified . 21 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Qwen Max scores higher in 1 category and Qwen3 Max in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 Max leads 48.1 to 30.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 16.1% for Qwen Max and 73.3% for Qwen3 Max.
  • Qwen3 Max is cheaper at $1.20 / $6 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Qwen3 Max accepts more context: 262K tokens versus 33K.

Side by side

Qwen Max and Qwen3 Max specifications
Qwen MaxQwen3 Max
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index34.743.7
Released2024-04-032025-09-23
WeightsProprietaryProprietary
Context window33K262K
Max output8K66K
Input $ / M tokens$1.60$1.20
Output $ / M tokens$6.40$6
Results tracked2333

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 Max leads

Qwen Max: 30.7 (#292), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkQwen MaxQwen3 Max
LMArena Coding12881456
Aider Polyglot21.8%—
ALE-Bench—370.45

Agentic & Tool Use Not comparable

Qwen Max: —, Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkQwen MaxQwen3 Max
Vending-Bench 2—71.56

Reasoning Qwen Max leads

Qwen Max: 25.1 (#151), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkQwen MaxQwen3 Max
LMArena Hard Prompts12691448
Kagi LLM Benchmark—72.5%
NYT Connections (extended)—30.1%
Chess Puzzles—4%
Mystery Game Puzzles—5%
DTBench—82.1%
LMCA—28.3%
Epoch Capabilities Index—142.38

Math Qwen3 Max leads

Qwen Max: 22.3 (#276), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkQwen MaxQwen3 Max
OTIS Mock AIME 2024-202516.1%73.3%
LMArena Math12751446
MATH Level 567.2%97.1%
FrontierMath (Tiers 1-3)—18.9%
FrontierMath (Feb 2025 set)1%—

Knowledge Qwen3 Max leads

Qwen Max: 30.3 (#228), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkQwen MaxQwen3 Max
GPQA Diamond56.1%72.6%
LMArena Expert12481455
SimpleQA Verified—48.7%

Multilingual Qwen3 Max leads

Qwen Max: 41.8 (#202), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkQwen MaxQwen3 Max
LMArena Non-English12631429
LMArena Chinese12541478
LMArena French13301449
LMArena German12541463
LMArena Japanese12051397
LMArena Korean11421399
LMArena Russian12741428
LMArena Spanish12901462

Instruction Following Qwen3 Max leads

Qwen Max: 66.5 (#208), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkQwen MaxQwen3 Max
LMArena Instruction Following12621419

Long Context Qwen3 Max leads

Qwen Max: 39.4 (#180), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkQwen MaxQwen3 Max
Fiction.LiveBench66.7%66.7%
LMArena Longer Query12881438
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

Qwen Max: 47.8 (#205), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkQwen MaxQwen3 Max
LMArena Text12821439
LMArena Creative Writing12481402
LMArena Multi-Turn12771446

Frequently asked questions

Is Qwen Max better than Qwen3 Max?

Qwen3 Max is the stronger model overall, scoring 43.7 to 34.7 on the Noometry Index.

Which is cheaper, Qwen Max or Qwen3 Max?

Qwen3 Max is cheaper. It lists at $1.20 per million input tokens and $6 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Qwen Max or Qwen3 Max better for coding?

Qwen3 Max scores higher on coding benchmarks: 43.0 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Qwen3 Max does, with 262K tokens against 33K.

How many benchmarks do Qwen Max and Qwen3 Max share?

21 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper