Model comparison

Claude 2.1 vs Qwen1.5 4b Chat

Qwen1.5 4b Chat is the stronger model overall, scoring 28.8 to 25.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • The widest gap is in math, where Qwen1.5 4b Chat leads 30.4 to 10.2.
  • Qwen1.5 4b Chat has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and Qwen1.5 4b Chat specifications
Claude 2.1Qwen1.5 4b Chat
ProviderAnthropicAlibaba (Qwen)
Noometry Index25.228.8
Released2023-11-21—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked713

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5 4b Chat leads

Claude 2.1: 26.2 (#327), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkClaude 2.1Qwen1.5 4b Chat
WeirdML7.1%—
LMArena Coding—999

Reasoning Claude 2.1 leads

Claude 2.1: 21.4 (#221), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkClaude 2.1Qwen1.5 4b Chat
LMArena Hard Prompts—976
DTBench51%—
Epoch Capabilities Index119.27—
ForecastBench54.2—

Math Qwen1.5 4b Chat leads

Claude 2.1: 10.2 (#315), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkClaude 2.1Qwen1.5 4b Chat
OTIS Mock AIME 2024-20251.9%—
LMArena Math—1026

Knowledge Qwen1.5 4b Chat leads

Claude 2.1: 15.4 (#292), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkClaude 2.1Qwen1.5 4b Chat
GPQA Diamond33%—
LMArena Expert—980
MMLU73.5%—

Multilingual Not comparable

Claude 2.1: —, Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkClaude 2.1Qwen1.5 4b Chat
LMArena Non-English—979
LMArena Chinese—1024
LMArena German—902
LMArena Russian—952

Instruction Following Not comparable

Claude 2.1: —, Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkClaude 2.1Qwen1.5 4b Chat
LMArena Instruction Following—978

Long Context Not comparable

Claude 2.1: —, Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkClaude 2.1Qwen1.5 4b Chat
LMArena Longer Query—988

Writing & Preference Not comparable

Claude 2.1: —, Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkClaude 2.1Qwen1.5 4b Chat
LMArena Text—997
LMArena Creative Writing—969
LMArena Multi-Turn—977

Frequently asked questions

Is Claude 2.1 better than Qwen1.5 4b Chat?

Qwen1.5 4b Chat is the stronger model overall, scoring 28.8 to 25.2 on the Noometry Index.

Is Claude 2.1 or Qwen1.5 4b Chat better for coding?

Qwen1.5 4b Chat scores higher on coding benchmarks: 29.1 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Qwen1.5 4b Chat share?

0 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper