Model comparison

Qwen3.7 Max vs Qwen3-Coder 480B-A35B Instruct

Qwen3.7 Max is the stronger model overall, scoring 51.5 to 38.1 on the Noometry Index.

Last verified . 14 shared benchmarks.

Qwen3.7 Max Alibaba (Qwen)

51.5

Rank #42 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Qwen3.7 Max scores higher in 8 categories and Qwen3-Coder 480B-A35B Instruct in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.7 Max leads 62.4 to 37.6.
  • Qwen3-Coder 480B-A35B Instruct is cheaper at $1.50 / $7.50 per million input/output tokens, against $2.50 / $7.50 for Qwen3.7 Max.
  • Qwen3.7 Max accepts more context: 1M tokens versus 262K.
  • Qwen3-Coder 480B-A35B Instruct has downloadable open weights; the other is API-only.

Side by side

Qwen3.7 Max and Qwen3-Coder 480B-A35B Instruct specifications
Qwen3.7 MaxQwen3-Coder 480B-A35B Instruct
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index51.538.1
Released2026-05-192025-04
WeightsProprietaryOpen
Context window1M262K
Max output131K66K
Input $ / M tokens$2.50$1.50
Output $ / M tokens$7.50$7.50
Results tracked3325

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.7 Max leads

Qwen3.7 Max: 50.4 (#45), Qwen3-Coder 480B-A35B Instruct: 35.5 (#223)

Coding benchmarks
BenchmarkQwen3.7 MaxQwen3-Coder 480B-A35B Instruct
LMArena WebDev15151275
LMArena Coding14981412
ALE-Bench1,189461.45
SWE-bench Verified77.3%—
SWE-bench Verified (bash only)—55.4%
SciCode48.8%—
GSO—4.9%
WeirdML—41.2%
AlgoTune—1.44

Agentic & Tool Use Qwen3-Coder 480B-A35B Instruct leads

Qwen3.7 Max: 22.1 (#135), Qwen3-Coder 480B-A35B Instruct: 23.9 (#123)

Agentic & Tool Use benchmarks
BenchmarkQwen3.7 MaxQwen3-Coder 480B-A35B Instruct
Terminal-Bench—27.2%
GBAEval0.4%—

Reasoning Qwen3.7 Max leads

Qwen3.7 Max: 49.2 (#38), Qwen3-Coder 480B-A35B Instruct: 25.5 (#149)

Reasoning benchmarks
BenchmarkQwen3.7 MaxQwen3-Coder 480B-A35B Instruct
LMArena Hard Prompts14831372
SimpleBench70.4%—
Kagi LLM Benchmark—49.5%
NYT Connections (extended)85.1%—
CritPt13.4%—
Chess Puzzles19%—
EBR-Bench9.5%—
Mystery Game Puzzles32%—
DTBench92.3%—
LMCA44%—
Epoch Capabilities Index153.68—

Math Qwen3.7 Max leads

Qwen3.7 Max: 62.4 (#32), Qwen3-Coder 480B-A35B Instruct: 37.6 (#150)

Math benchmarks
BenchmarkQwen3.7 MaxQwen3-Coder 480B-A35B Instruct
LMArena Math14901365
FrontierMath (Tiers 1-3)64.6%—
FrontierMath Tier 434.1%—
OTIS Mock AIME 2024-202595.6%—
ProofBench26%—

Knowledge Qwen3.7 Max leads

Qwen3.7 Max: 61.6 (#28), Qwen3-Coder 480B-A35B Instruct: 37.0 (#162)

Knowledge benchmarks
BenchmarkQwen3.7 MaxQwen3-Coder 480B-A35B Instruct
LMArena Expert14881338
GPQA Diamond90.9%—
SimpleQA Verified55.8%—

Multilingual Qwen3.7 Max leads

Qwen3.7 Max: 56.9 (#15), Qwen3-Coder 480B-A35B Instruct: 47.7 (#148)

Multilingual benchmarks
BenchmarkQwen3.7 MaxQwen3-Coder 480B-A35B Instruct
LMArena Non-English14741346
LMArena Chinese15301357
LMArena Russian14841366
LMArena French—1398
LMArena German—1325
LMArena Japanese—1310
LMArena Korean—1305
LMArena Spanish—1360

Instruction Following Qwen3.7 Max leads

Qwen3.7 Max: 76.7 (#38), Qwen3-Coder 480B-A35B Instruct: 71.6 (#147)

Instruction Following benchmarks
BenchmarkQwen3.7 MaxQwen3-Coder 480B-A35B Instruct
LMArena Instruction Following14601355

Long Context Qwen3.7 Max leads

Qwen3.7 Max: 45.4 (#40), Qwen3-Coder 480B-A35B Instruct: 42.0 (#131)

Long Context benchmarks
BenchmarkQwen3.7 MaxQwen3-Coder 480B-A35B Instruct
LMArena Longer Query14821378

Writing & Preference Qwen3.7 Max leads

Qwen3.7 Max: 65.0 (#54), Qwen3-Coder 480B-A35B Instruct: 55.3 (#147)

Writing & Preference benchmarks
BenchmarkQwen3.7 MaxQwen3-Coder 480B-A35B Instruct
LMArena Text14761357
LMArena Creative Writing14491333
LMArena Multi-Turn14811365
EQ-Bench 41110—

Frequently asked questions

Is Qwen3.7 Max better than Qwen3-Coder 480B-A35B Instruct?

Qwen3.7 Max is the stronger model overall, scoring 51.5 to 38.1 on the Noometry Index.

Which is cheaper, Qwen3.7 Max or Qwen3-Coder 480B-A35B Instruct?

Qwen3-Coder 480B-A35B Instruct is cheaper. It lists at $1.50 per million input tokens and $7.50 per million output tokens; Qwen3.7 Max lists at $2.50 and $7.50.

Is Qwen3.7 Max or Qwen3-Coder 480B-A35B Instruct better for coding?

Qwen3.7 Max scores higher on coding benchmarks: 50.4 versus 35.5 in the Noometry coding category.

Which has the bigger context window?

Qwen3.7 Max does, with 1M tokens against 262K.

How many benchmarks do Qwen3.7 Max and Qwen3-Coder 480B-A35B Instruct share?

14 benchmarks have published results for both models. Qwen3.7 Max has 33 scored results on Noometry and Qwen3-Coder 480B-A35B Instruct has 25.

Related comparisons

Go deeper