Model comparison

Phi 3 Mini 128k Instruct vs Qwen-14B

Qwen-14B is the stronger model overall, scoring 31.4 to 29.7 on the Noometry Index.

Last verified . 10 shared benchmarks.

Phi 3 Mini 128k Instruct Microsoft

29.7

Rank #305 Confirmed

Qwen-14B Alibaba (Qwen)

31.4

Rank #275 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Phi 3 Mini 128k Instruct scores higher in 2 categories and Qwen-14B in 5 categories; 2 gaps are clear of the uncertainty.

Side by side

Phi 3 Mini 128k Instruct and Qwen-14B specifications
Phi 3 Mini 128k InstructQwen-14B
ProviderMicrosoftAlibaba (Qwen)
Noometry Index29.731.4
Released2024-04-232023-09-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1918

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen-14B leads

Phi 3 Mini 128k Instruct: 28.8 (#312), Qwen-14B: 31.2 (#288)

Coding benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen-14B
LMArena Coding10391071
BigCodeBench Instruct29.6%—
BigCodeBench Complete40.6%—

Reasoning Too close to call

Phi 3 Mini 128k Instruct: 19.6 (#256), Qwen-14B: 19.6 (#257)

Reasoning benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen-14B
LMArena Hard Prompts10281027
BIG-Bench Hard—55%
Epoch Capabilities Index—113.03
LAMBADA—71.1%
PIQA—79.9%

Math Too close to call

Phi 3 Mini 128k Instruct: 31.6 (#222), Qwen-14B: 31.2 (#227)

Math benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen-14B
LMArena Math10891068
GSM8K—61.3%

Knowledge Not comparable

Phi 3 Mini 128k Instruct: 26.8 (#254), Qwen-14B: —

Knowledge benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen-14B
LMArena Expert984—
ARC (AI2) Challenge—84.4%
BoolQ—86.2%
MMLU—66.3%

Multilingual Qwen-14B leads

Phi 3 Mini 128k Instruct: 25.2 (#285), Qwen-14B: 27.5 (#275)

Multilingual benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen-14B
LMArena Non-English10001041
LMArena Chinese10161077
LMArena French1039—
LMArena German1006—
LMArena Japanese899—
LMArena Korean856—
LMArena Russian1004—
LMArena Spanish1059—

Instruction Following Too close to call

Phi 3 Mini 128k Instruct: 51.8 (#294), Qwen-14B: 52.4 (#289)

Instruction Following benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen-14B
LMArena Instruction Following10231031

Long Context Too close to call

Phi 3 Mini 128k Instruct: 30.4 (#289), Qwen-14B: 31.3 (#280)

Long Context benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen-14B
LMArena Longer Query9961028

Writing & Preference Too close to call

Phi 3 Mini 128k Instruct: 27.1 (#301), Qwen-14B: 27.6 (#299)

Writing & Preference benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen-14B
LMArena Text10501051
LMArena Creative Writing10241028
LMArena Multi-Turn9891022

Frequently asked questions

Is Phi 3 Mini 128k Instruct better than Qwen-14B?

Qwen-14B is the stronger model overall, scoring 31.4 to 29.7 on the Noometry Index.

Is Phi 3 Mini 128k Instruct or Qwen-14B better for coding?

Qwen-14B scores higher on coding benchmarks: 31.2 versus 28.8 in the Noometry coding category.

How many benchmarks do Phi 3 Mini 128k Instruct and Qwen-14B share?

10 benchmarks have published results for both models. Phi 3 Mini 128k Instruct has 19 scored results on Noometry and Qwen-14B has 18.

Related comparisons

Go deeper