Model comparison

Qwen3.7 Max vs Step 1o Turbo 202506

Qwen3.7 Max is the stronger model overall, scoring 51.5 to 39.7 on the Noometry Index.

Last verified . 12 shared benchmarks.

Qwen3.7 Max Alibaba (Qwen)

51.5

Rank #42 Confirmed

Step 1o Turbo 202506 StepFun

39.7

Rank #160 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Qwen3.7 Max scores higher in 8 categories and Step 1o Turbo 202506 in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.7 Max leads 62.4 to 36.6.

Side by side

Qwen3.7 Max and Step 1o Turbo 202506 specifications
Qwen3.7 MaxStep 1o Turbo 202506
ProviderAlibaba (Qwen)StepFun
Noometry Index51.539.7
Released2026-05-19—
WeightsProprietaryProprietary
Context window1M—
Max output131K—
Input $ / M tokens$2.50—
Output $ / M tokens$7.50—
Results tracked3314

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.7 Max leads

Qwen3.7 Max: 50.4 (#45), Step 1o Turbo 202506: 39.2 (#160)

Coding benchmarks
BenchmarkQwen3.7 MaxStep 1o Turbo 202506
LMArena Coding14981339
SWE-bench Verified77.3%—
LMArena WebDev1515—
SciCode48.8%—
ALE-Bench1,189—

Agentic & Tool Use Not comparable

Qwen3.7 Max: 22.1 (#135), Step 1o Turbo 202506: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.7 MaxStep 1o Turbo 202506
GBAEval0.4%—

Reasoning Qwen3.7 Max leads

Qwen3.7 Max: 49.2 (#38), Step 1o Turbo 202506: 26.8 (#129)

Reasoning benchmarks
BenchmarkQwen3.7 MaxStep 1o Turbo 202506
LMArena Hard Prompts14831335
SimpleBench70.4%—
NYT Connections (extended)85.1%—
CritPt13.4%—
Chess Puzzles19%—
EBR-Bench9.5%—
Mystery Game Puzzles32%—
DTBench92.3%—
LMCA44%—
Epoch Capabilities Index153.68—

Math Qwen3.7 Max leads

Qwen3.7 Max: 62.4 (#32), Step 1o Turbo 202506: 36.6 (#164)

Math benchmarks
BenchmarkQwen3.7 MaxStep 1o Turbo 202506
LMArena Math14901318
FrontierMath (Tiers 1-3)64.6%—
FrontierMath Tier 434.1%—
OTIS Mock AIME 2024-202595.6%—
ProofBench26%—

Knowledge Qwen3.7 Max leads

Qwen3.7 Max: 61.6 (#28), Step 1o Turbo 202506: 36.1 (#176)

Knowledge benchmarks
BenchmarkQwen3.7 MaxStep 1o Turbo 202506
LMArena Expert14881308
GPQA Diamond90.9%—
SimpleQA Verified55.8%—

Multimodal Not comparable

Qwen3.7 Max: —, Step 1o Turbo 202506: 36.1 (#80)

Multimodal benchmarks
BenchmarkQwen3.7 MaxStep 1o Turbo 202506
LMArena Vision—1186

Multilingual Qwen3.7 Max leads

Qwen3.7 Max: 56.9 (#15), Step 1o Turbo 202506: 45.3 (#173)

Multilingual benchmarks
BenchmarkQwen3.7 MaxStep 1o Turbo 202506
LMArena Non-English14741313
LMArena Chinese15301380
LMArena Russian14841327
LMArena German—1314

Instruction Following Qwen3.7 Max leads

Qwen3.7 Max: 76.7 (#38), Step 1o Turbo 202506: 69.2 (#176)

Instruction Following benchmarks
BenchmarkQwen3.7 MaxStep 1o Turbo 202506
LMArena Instruction Following14601310

Long Context Qwen3.7 Max leads

Qwen3.7 Max: 45.4 (#40), Step 1o Turbo 202506: 41.0 (#146)

Long Context benchmarks
BenchmarkQwen3.7 MaxStep 1o Turbo 202506
LMArena Longer Query14821348

Writing & Preference Qwen3.7 Max leads

Qwen3.7 Max: 65.0 (#54), Step 1o Turbo 202506: 53.1 (#160)

Writing & Preference benchmarks
BenchmarkQwen3.7 MaxStep 1o Turbo 202506
LMArena Text14761336
LMArena Creative Writing14491307
LMArena Multi-Turn14811340
EQ-Bench 41110—

Frequently asked questions

Is Qwen3.7 Max better than Step 1o Turbo 202506?

Qwen3.7 Max is the stronger model overall, scoring 51.5 to 39.7 on the Noometry Index.

Is Qwen3.7 Max or Step 1o Turbo 202506 better for coding?

Qwen3.7 Max scores higher on coding benchmarks: 50.4 versus 39.2 in the Noometry coding category.

How many benchmarks do Qwen3.7 Max and Step 1o Turbo 202506 share?

12 benchmarks have published results for both models. Qwen3.7 Max has 33 scored results on Noometry and Step 1o Turbo 202506 has 14.

Related comparisons

Go deeper