Model comparison

Phi 3 Small 8k Instruct vs Qwen1.5 4b Chat

Phi 3 Small 8k Instruct and Qwen1.5 4b Chat score almost the same on the Noometry Index (29.3 vs 28.8), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Phi 3 Small 8k Instruct Microsoft

29.3

Rank #314 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Phi 3 Small 8k Instruct scores higher in 5 categories and Qwen1.5 4b Chat in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Phi 3 Small 8k Instruct leads 31.1 to 23.8.

Side by side

Phi 3 Small 8k Instruct and Qwen1.5 4b Chat specifications
Phi 3 Small 8k InstructQwen1.5 4b Chat
ProviderMicrosoftAlibaba (Qwen)
Noometry Index29.328.8
Released2024-04-23—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3213

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5 4b Chat leads

Phi 3 Small 8k Instruct: 27.9 (#318), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkPhi 3 Small 8k InstructQwen1.5 4b Chat
LMArena Coding1101999
LiveBench Coding20.3%—

Reasoning Qwen1.5 4b Chat leads

Phi 3 Small 8k Instruct: 14.8 (#323), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkPhi 3 Small 8k InstructQwen1.5 4b Chat
LMArena Hard Prompts1100976
LiveBench Reasoning15.9%—
LiveBench Data Analysis30.3%—
Adversarial NLI58.1%—
BIG-Bench Hard79.1%—
HellaSwag77%—
LiveBench24%—
WinoGrande81.5%—

Math Qwen1.5 4b Chat leads

Phi 3 Small 8k Instruct: 27.6 (#248), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkPhi 3 Small 8k InstructQwen1.5 4b Chat
LMArena Math11511026
LiveBench Math17.6%—

Knowledge Phi 3 Small 8k Instruct leads

Phi 3 Small 8k Instruct: 29.1 (#240), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkPhi 3 Small 8k InstructQwen1.5 4b Chat
LMArena Expert1067980
ARC (AI2) Challenge90.7%—
MMLU75.7%—
OpenBookQA88%—
TriviaQA58.1%—

Multilingual Phi 3 Small 8k Instruct leads

Phi 3 Small 8k Instruct: 28.5 (#272), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkPhi 3 Small 8k InstructQwen1.5 4b Chat
LMArena Non-English1058979
LMArena Chinese10611024
LMArena German1080902
LMArena Russian1111952
LMArena French1135—
LMArena Japanese966—
LMArena Korean894—
LMArena Spanish1111—

Instruction Following Phi 3 Small 8k Instruct leads

Phi 3 Small 8k Instruct: 51.9 (#292), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkPhi 3 Small 8k InstructQwen1.5 4b Chat
LMArena Instruction Following1087978
LiveBench Instruction Following47.2%—

Long Context Phi 3 Small 8k Instruct leads

Phi 3 Small 8k Instruct: 33.0 (#267), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkPhi 3 Small 8k InstructQwen1.5 4b Chat
LMArena Longer Query1088988

Writing & Preference Phi 3 Small 8k Instruct leads

Phi 3 Small 8k Instruct: 31.1 (#284), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkPhi 3 Small 8k InstructQwen1.5 4b Chat
LMArena Text1110997
LMArena Creative Writing1083969
LMArena Multi-Turn1068977
LiveBench Language12.9%—

Frequently asked questions

Is Phi 3 Small 8k Instruct better than Qwen1.5 4b Chat?

Phi 3 Small 8k Instruct and Qwen1.5 4b Chat score almost the same on the Noometry Index (29.3 vs 28.8), so choose on price, context window or the category you care about most.

Is Phi 3 Small 8k Instruct or Qwen1.5 4b Chat better for coding?

Qwen1.5 4b Chat scores higher on coding benchmarks: 29.1 versus 27.9 in the Noometry coding category.

How many benchmarks do Phi 3 Small 8k Instruct and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Phi 3 Small 8k Instruct has 32 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper