Model comparison

Phi 3 Mini 128k Instruct vs Qwen1.5 4b Chat

Phi 3 Mini 128k Instruct and Qwen1.5 4b Chat score almost the same on the Noometry Index (29.7 vs 28.8), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Phi 3 Mini 128k Instruct Microsoft

29.7

Rank #305 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Phi 3 Mini 128k Instruct scores higher in 7 categories and Qwen1.5 4b Chat in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Phi 3 Mini 128k Instruct leads 27.1 to 23.8.

Side by side

Phi 3 Mini 128k Instruct and Qwen1.5 4b Chat specifications
Phi 3 Mini 128k InstructQwen1.5 4b Chat
ProviderMicrosoftAlibaba (Qwen)
Noometry Index29.728.8
Released2024-04-23—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1913

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Phi 3 Mini 128k Instruct: 28.8 (#312), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5 4b Chat
LMArena Coding1039999
BigCodeBench Instruct29.6%—
BigCodeBench Complete40.6%—

Reasoning Phi 3 Mini 128k Instruct leads

Phi 3 Mini 128k Instruct: 19.6 (#256), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5 4b Chat
LMArena Hard Prompts1028976

Math Phi 3 Mini 128k Instruct leads

Phi 3 Mini 128k Instruct: 31.6 (#222), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5 4b Chat
LMArena Math10891026

Knowledge Too close to call

Phi 3 Mini 128k Instruct: 26.8 (#254), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5 4b Chat
LMArena Expert984980

Multilingual Phi 3 Mini 128k Instruct leads

Phi 3 Mini 128k Instruct: 25.2 (#285), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5 4b Chat
LMArena Non-English1000979
LMArena Chinese10161024
LMArena German1006902
LMArena Russian1004952
LMArena French1039—
LMArena Japanese899—
LMArena Korean856—
LMArena Spanish1059—

Instruction Following Phi 3 Mini 128k Instruct leads

Phi 3 Mini 128k Instruct: 51.8 (#294), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5 4b Chat
LMArena Instruction Following1023978

Long Context Too close to call

Phi 3 Mini 128k Instruct: 30.4 (#289), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5 4b Chat
LMArena Longer Query996988

Writing & Preference Phi 3 Mini 128k Instruct leads

Phi 3 Mini 128k Instruct: 27.1 (#301), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5 4b Chat
LMArena Text1050997
LMArena Creative Writing1024969
LMArena Multi-Turn989977

Frequently asked questions

Is Phi 3 Mini 128k Instruct better than Qwen1.5 4b Chat?

Phi 3 Mini 128k Instruct and Qwen1.5 4b Chat score almost the same on the Noometry Index (29.7 vs 28.8), so choose on price, context window or the category you care about most.

Is Phi 3 Mini 128k Instruct or Qwen1.5 4b Chat better for coding?

They score almost the same on coding (28.8 vs 29.1); test both on your own repository before choosing.

How many benchmarks do Phi 3 Mini 128k Instruct and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Phi 3 Mini 128k Instruct has 19 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper