Model comparison

Qwen1.5 4b Chat vs Qwen3.7 Flash

Qwen3.7 Flash is the stronger model overall, scoring 39.9 to 28.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Qwen3.7 Flash Alibaba (Qwen)

39.9

Rank #156 Confirmed

Summary

  • The widest gap is in knowledge, where Qwen3.7 Flash leads 48.9 to 26.7.
  • Qwen1.5 4b Chat has downloadable open weights; the other is API-only.

Side by side

Qwen1.5 4b Chat and Qwen3.7 Flash specifications
Qwen1.5 4b ChatQwen3.7 Flash
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index28.839.9
Released—2026-07-15
WeightsOpenProprietary
Context window—1M
Max output—131K
Input $ / M tokens—$0.03
Output $ / M tokens—$0.13
Results tracked137

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Qwen1.5 4b Chat: 29.1 (#308), Qwen3.7 Flash: —

Coding benchmarks
BenchmarkQwen1.5 4b ChatQwen3.7 Flash
LMArena Coding999—

Reasoning Qwen3.7 Flash leads

Qwen1.5 4b Chat: 18.5 (#279), Qwen3.7 Flash: 28.2 (#108)

Reasoning benchmarks
BenchmarkQwen1.5 4b ChatQwen3.7 Flash
NYT Connections (extended)—43.8%
Chess Puzzles—23%
LMArena Hard Prompts976—
Mystery Game Puzzles—15%
Epoch Capabilities Index—144.64

Math Qwen3.7 Flash leads

Qwen1.5 4b Chat: 30.4 (#234), Qwen3.7 Flash: 38.3 (#140)

Math benchmarks
BenchmarkQwen1.5 4b ChatQwen3.7 Flash
FrontierMath (Tiers 1-3)—19.3%
OTIS Mock AIME 2024-2025—86.7%
LMArena Math1026—

Knowledge Qwen3.7 Flash leads

Qwen1.5 4b Chat: 26.7 (#255), Qwen3.7 Flash: 48.9 (#75)

Knowledge benchmarks
BenchmarkQwen1.5 4b ChatQwen3.7 Flash
GPQA Diamond—82.3%
LMArena Expert980—

Multilingual Not comparable

Qwen1.5 4b Chat: 24.1 (#290), Qwen3.7 Flash: —

Multilingual benchmarks
BenchmarkQwen1.5 4b ChatQwen3.7 Flash
LMArena Non-English979—
LMArena Chinese1024—
LMArena German902—
LMArena Russian952—

Instruction Following Not comparable

Qwen1.5 4b Chat: 49.0 (#300), Qwen3.7 Flash: —

Instruction Following benchmarks
BenchmarkQwen1.5 4b ChatQwen3.7 Flash
LMArena Instruction Following978—

Long Context Not comparable

Qwen1.5 4b Chat: 30.1 (#290), Qwen3.7 Flash: —

Long Context benchmarks
BenchmarkQwen1.5 4b ChatQwen3.7 Flash
LMArena Longer Query988—

Writing & Preference Not comparable

Qwen1.5 4b Chat: 23.8 (#309), Qwen3.7 Flash: —

Writing & Preference benchmarks
BenchmarkQwen1.5 4b ChatQwen3.7 Flash
LMArena Text997—
LMArena Creative Writing969—
LMArena Multi-Turn977—

Frequently asked questions

Is Qwen1.5 4b Chat better than Qwen3.7 Flash?

Qwen3.7 Flash is the stronger model overall, scoring 39.9 to 28.8 on the Noometry Index.

How many benchmarks do Qwen1.5 4b Chat and Qwen3.7 Flash share?

0 benchmarks have published results for both models. Qwen1.5 4b Chat has 13 scored results on Noometry and Qwen3.7 Flash has 7.

Related comparisons

Go deeper