Model comparison

Qwen Max vs Qwen1.5 4b Chat

Qwen Max is the stronger model overall, scoring 34.7 to 28.8 on the Noometry Index.

Last verified . 13 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Qwen Max scores higher in 7 categories and Qwen1.5 4b Chat in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Max leads 47.8 to 23.8.
  • Qwen1.5 4b Chat has downloadable open weights; the other is API-only.

Side by side

Qwen Max and Qwen1.5 4b Chat specifications
Qwen MaxQwen1.5 4b Chat
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index34.728.8
Released2024-04-03—
WeightsProprietaryOpen
Context window33K—
Max output8K—
Input $ / M tokens$1.60—
Output $ / M tokens$6.40—
Results tracked2313

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Max leads

Qwen Max: 30.7 (#292), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkQwen MaxQwen1.5 4b Chat
LMArena Coding1288999
Aider Polyglot21.8%—

Reasoning Qwen Max leads

Qwen Max: 25.1 (#151), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkQwen MaxQwen1.5 4b Chat
LMArena Hard Prompts1269976

Math Qwen1.5 4b Chat leads

Qwen Max: 22.3 (#276), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkQwen MaxQwen1.5 4b Chat
LMArena Math12751026
OTIS Mock AIME 2024-202516.1%—
MATH Level 567.2%—
FrontierMath (Feb 2025 set)1%—

Knowledge Qwen Max leads

Qwen Max: 30.3 (#228), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkQwen MaxQwen1.5 4b Chat
LMArena Expert1248980
GPQA Diamond56.1%—

Multilingual Qwen Max leads

Qwen Max: 41.8 (#202), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkQwen MaxQwen1.5 4b Chat
LMArena Non-English1263979
LMArena Chinese12541024
LMArena German1254902
LMArena Russian1274952
LMArena French1330—
LMArena Japanese1205—
LMArena Korean1142—
LMArena Spanish1290—

Instruction Following Qwen Max leads

Qwen Max: 66.5 (#208), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkQwen MaxQwen1.5 4b Chat
LMArena Instruction Following1262978

Long Context Qwen Max leads

Qwen Max: 39.4 (#180), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkQwen MaxQwen1.5 4b Chat
LMArena Longer Query1288988
Fiction.LiveBench66.7%—

Writing & Preference Qwen Max leads

Qwen Max: 47.8 (#205), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkQwen MaxQwen1.5 4b Chat
LMArena Text1282997
LMArena Creative Writing1248969
LMArena Multi-Turn1277977

Frequently asked questions

Is Qwen Max better than Qwen1.5 4b Chat?

Qwen Max is the stronger model overall, scoring 34.7 to 28.8 on the Noometry Index.

Is Qwen Max or Qwen1.5 4b Chat better for coding?

Qwen Max scores higher on coding benchmarks: 30.7 versus 29.1 in the Noometry coding category.

How many benchmarks do Qwen Max and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper