Model comparison

Phi 3 Mini 128k Instruct vs Qwen1.5-32B

Phi 3 Mini 128k Instruct and Qwen1.5-32B score almost the same on the Noometry Index (29.7 vs 30.5), so choose on price, context window or the category you care about most.

Last verified . 19 shared benchmarks.

Phi 3 Mini 128k Instruct Microsoft

29.7

Rank #305 Confirmed

Qwen1.5-32B Alibaba (Qwen)

30.5

Rank #293 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Phi 3 Mini 128k Instruct scores higher in 1 category and Qwen1.5-32B in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Phi 3 Mini 128k Instruct leads 26.8 to 13.5.

Side by side

Phi 3 Mini 128k Instruct and Qwen1.5-32B specifications
Phi 3 Mini 128k InstructQwen1.5-32B
ProviderMicrosoftAlibaba (Qwen)
Noometry Index29.730.5
Released2024-04-232024-02-04
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1921

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5-32B leads

Phi 3 Mini 128k Instruct: 28.8 (#312), Qwen1.5-32B: 31.7 (#282)

Coding benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5-32B
BigCodeBench Instruct29.6%32.3%
LMArena Coding10391155
BigCodeBench Complete40.6%42%

Reasoning Qwen1.5-32B leads

Phi 3 Mini 128k Instruct: 19.6 (#256), Qwen1.5-32B: 21.8 (#212)

Reasoning benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5-32B
LMArena Hard Prompts10281130

Math Qwen1.5-32B leads

Phi 3 Mini 128k Instruct: 31.6 (#222), Qwen1.5-32B: 33.0 (#207)

Math benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5-32B
LMArena Math10891155

Knowledge Phi 3 Mini 128k Instruct leads

Phi 3 Mini 128k Instruct: 26.8 (#254), Qwen1.5-32B: 13.5 (#296)

Knowledge benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5-32B
LMArena Expert9841126
GPQA Diamond—30.7%
MMLU—74.4%

Multilingual Qwen1.5-32B leads

Phi 3 Mini 128k Instruct: 25.2 (#285), Qwen1.5-32B: 31.4 (#259)

Multilingual benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5-32B
LMArena Non-English10001106
LMArena Chinese10161177
LMArena French10391101
LMArena German10061058
LMArena Japanese8991027
LMArena Korean8561008
LMArena Russian10041073
LMArena Spanish10591089

Instruction Following Qwen1.5-32B leads

Phi 3 Mini 128k Instruct: 51.8 (#294), Qwen1.5-32B: 57.7 (#265)

Instruction Following benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5-32B
LMArena Instruction Following10231116

Long Context Qwen1.5-32B leads

Phi 3 Mini 128k Instruct: 30.4 (#289), Qwen1.5-32B: 34.7 (#246)

Long Context benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5-32B
LMArena Longer Query9961146

Writing & Preference Qwen1.5-32B leads

Phi 3 Mini 128k Instruct: 27.1 (#301), Qwen1.5-32B: 34.2 (#271)

Writing & Preference benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen1.5-32B
LMArena Text10501137
LMArena Creative Writing10241083
LMArena Multi-Turn9891140

Frequently asked questions

Is Phi 3 Mini 128k Instruct better than Qwen1.5-32B?

Phi 3 Mini 128k Instruct and Qwen1.5-32B score almost the same on the Noometry Index (29.7 vs 30.5), so choose on price, context window or the category you care about most.

Is Phi 3 Mini 128k Instruct or Qwen1.5-32B better for coding?

Qwen1.5-32B scores higher on coding benchmarks: 31.7 versus 28.8 in the Noometry coding category.

How many benchmarks do Phi 3 Mini 128k Instruct and Qwen1.5-32B share?

19 benchmarks have published results for both models. Phi 3 Mini 128k Instruct has 19 scored results on Noometry and Qwen1.5-32B has 21.

Related comparisons

Go deeper