Model comparison

Qwen3.5-Flash vs Qwen3.7 Plus

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 42.5 on the Noometry Index. Qwen3.5-Flash costs 4.0× less per token, which makes it the better buy when Qwen3.7 Plus's lead doesn't matter for your workload.

Last verified . 25 shared benchmarks.

Qwen3.5-Flash Alibaba (Qwen)

42.5

Rank #112 Confirmed

Qwen3.7 Plus Alibaba (Qwen)

45.3

Rank #72 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Qwen3.5-Flash scores higher in 0 categories and Qwen3.7 Plus in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.7 Plus leads 50.5 to 37.4.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 18.2% for Qwen3.5-Flash and 34.4% for Qwen3.7 Plus.
  • Qwen3.5-Flash is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.40 / $1.60 for Qwen3.7 Plus.

Side by side

Qwen3.5-Flash and Qwen3.7 Plus specifications
Qwen3.5-FlashQwen3.7 Plus
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index42.545.3
Released2026-02-232026-06-02
WeightsProprietaryProprietary
Context window1M1M
Max output66K131K
Input $ / M tokens$0.10$0.40
Output $ / M tokens$0.40$1.60
Results tracked3232

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.7 Plus leads

Qwen3.5-Flash: 34.2 (#242), Qwen3.7 Plus: 36.6 (#206)

Coding benchmarks
BenchmarkQwen3.5-FlashQwen3.7 Plus
LMArena Coding14121473
FrontierCode—10.2%
LMArena WebDev1244—
SciCode—45.5%
ALE-Bench221.8—

Agentic & Tool Use Not comparable

Qwen3.5-Flash: —, Qwen3.7 Plus: 21.4 (#138)

Agentic & Tool Use benchmarks
BenchmarkQwen3.5-FlashQwen3.7 Plus
OSWorld 2.0—2.8%
Vending-Bench 2462.69—

Reasoning Qwen3.7 Plus leads

Qwen3.5-Flash: 33.7 (#72), Qwen3.7 Plus: 39.3 (#59)

Reasoning benchmarks
BenchmarkQwen3.5-FlashQwen3.7 Plus
Chess Puzzles21%24%
LMArena Hard Prompts14031460
Mystery Game Puzzles20%17%
DTBench82.9%84%
LMCA29.1%37.6%
Epoch Capabilities Index143.98147.37
NYT Connections (extended)—74.8%
CritPt—9.1%

Math Qwen3.7 Plus leads

Qwen3.5-Flash: 37.4 (#158), Qwen3.7 Plus: 50.5 (#56)

Math benchmarks
BenchmarkQwen3.5-FlashQwen3.7 Plus
FrontierMath (Tiers 1-3)18.2%34.4%
OTIS Mock AIME 2024-202584.4%93.3%
LMArena Math14071466
FrontierMath (Feb 2025 set)6.2%—
FrontierMath Tier 4 (v1)0%—

Knowledge Qwen3.7 Plus leads

Qwen3.5-Flash: 43.2 (#93), Qwen3.7 Plus: 54.9 (#51)

Knowledge benchmarks
BenchmarkQwen3.5-FlashQwen3.7 Plus
GPQA Diamond82.3%87.9%
LMArena Expert14071467
SimpleQA Verified20.3%—
Vectara Hallucination Rate10.5%—

Multimodal Not comparable

Qwen3.5-Flash: —, Qwen3.7 Plus: 41.8 (#33)

Multimodal benchmarks
BenchmarkQwen3.5-FlashQwen3.7 Plus
LMArena Vision—1279
LMArena Document—1444

Multilingual Qwen3.7 Plus leads

Qwen3.5-Flash: 50.5 (#121), Qwen3.7 Plus: 54.8 (#38)

Multilingual benchmarks
BenchmarkQwen3.5-FlashQwen3.7 Plus
LMArena Non-English13851445
LMArena Chinese14461510
LMArena French14121473
LMArena German13901471
LMArena Japanese13681413
LMArena Korean13441415
LMArena Russian13791457
LMArena Spanish14001457

Instruction Following Qwen3.7 Plus leads

Qwen3.5-Flash: 72.6 (#139), Qwen3.7 Plus: 75.8 (#52)

Instruction Following benchmarks
BenchmarkQwen3.5-FlashQwen3.7 Plus
LMArena Instruction Following13741440

Long Context Qwen3.7 Plus leads

Qwen3.5-Flash: 42.4 (#124), Qwen3.7 Plus: 44.5 (#65)

Long Context benchmarks
BenchmarkQwen3.5-FlashQwen3.7 Plus
LMArena Longer Query13921455

Writing & Preference Qwen3.7 Plus leads

Qwen3.5-Flash: 57.9 (#122), Qwen3.7 Plus: 64.3 (#56)

Writing & Preference benchmarks
BenchmarkQwen3.5-FlashQwen3.7 Plus
LMArena Text13971455
LMArena Creative Writing13431439
LMArena Multi-Turn13931460

Frequently asked questions

Is Qwen3.5-Flash better than Qwen3.7 Plus?

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 42.5 on the Noometry Index. Qwen3.5-Flash costs 4.0× less per token, which makes it the better buy when Qwen3.7 Plus's lead doesn't matter for your workload.

Which is cheaper, Qwen3.5-Flash or Qwen3.7 Plus?

Qwen3.5-Flash is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; Qwen3.7 Plus lists at $0.40 and $1.60.

Is Qwen3.5-Flash or Qwen3.7 Plus better for coding?

Qwen3.7 Plus scores higher on coding benchmarks: 36.6 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do Qwen3.5-Flash and Qwen3.7 Plus share?

25 benchmarks have published results for both models. Qwen3.5-Flash has 32 scored results on Noometry and Qwen3.7 Plus has 32.

Related comparisons

Go deeper