Model comparison

Phi 3 Mini 128k Instruct vs Qwen2-72B

Phi 3 Mini 128k Instruct and Qwen2-72B score almost the same on the Noometry Index (29.7 vs 30.0), so choose on price, context window or the category you care about most.

Last verified . 19 shared benchmarks.

Phi 3 Mini 128k Instruct Microsoft

29.7

Rank #305 Confirmed

Qwen2-72B Alibaba (Qwen)

30.0

Rank #300 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Phi 3 Mini 128k Instruct scores higher in 2 categories and Qwen2-72B in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2-72B leads 40.8 to 27.1.
  • The biggest single-benchmark swing is BigCodeBench Complete: 40.6% for Phi 3 Mini 128k Instruct and 54% for Qwen2-72B.

Side by side

Phi 3 Mini 128k Instruct and Qwen2-72B specifications
Phi 3 Mini 128k InstructQwen2-72B
ProviderMicrosoftAlibaba (Qwen)
Noometry Index29.730.0
Released2024-04-232024-06-07
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1926

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Phi 3 Mini 128k Instruct: 28.8 (#312), Qwen2-72B: 29.1 (#310)

Coding benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen2-72B
BigCodeBench Instruct29.6%38.5%
LMArena Coding10391196
BigCodeBench Complete40.6%54%
WeirdML—11.3%

Agentic & Tool Use Not comparable

Phi 3 Mini 128k Instruct: —, Qwen2-72B: 17.0 (#146)

Agentic & Tool Use benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen2-72B
TheAgentCompany—1.1%
METR Time Horizons—29.9%

Reasoning Qwen2-72B leads

Phi 3 Mini 128k Instruct: 19.6 (#256), Qwen2-72B: 23.2 (#181)

Reasoning benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen2-72B
LMArena Hard Prompts10281191
Epoch Capabilities Index—125.28

Math Phi 3 Mini 128k Instruct leads

Phi 3 Mini 128k Instruct: 31.6 (#222), Qwen2-72B: 30.2 (#236)

Math benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen2-72B
LMArena Math10891235
MATH Level 5—39.1%

Knowledge Phi 3 Mini 128k Instruct leads

Phi 3 Mini 128k Instruct: 26.8 (#254), Qwen2-72B: 21.2 (#275)

Knowledge benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen2-72B
LMArena Expert9841171
GPQA Diamond—40.8%
MMLU—82.4%

Multilingual Qwen2-72B leads

Phi 3 Mini 128k Instruct: 25.2 (#285), Qwen2-72B: 35.9 (#244)

Multilingual benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen2-72B
LMArena Non-English10001176
LMArena Chinese10161240
LMArena French10391170
LMArena German10061151
LMArena Japanese8991111
LMArena Korean8561083
LMArena Russian10041169
LMArena Spanish10591169

Instruction Following Qwen2-72B leads

Phi 3 Mini 128k Instruct: 51.8 (#294), Qwen2-72B: 61.7 (#241)

Instruction Following benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen2-72B
LMArena Instruction Following10231181

Long Context Qwen2-72B leads

Phi 3 Mini 128k Instruct: 30.4 (#289), Qwen2-72B: 36.1 (#235)

Long Context benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen2-72B
LMArena Longer Query9961192

Writing & Preference Qwen2-72B leads

Phi 3 Mini 128k Instruct: 27.1 (#301), Qwen2-72B: 40.8 (#241)

Writing & Preference benchmarks
BenchmarkPhi 3 Mini 128k InstructQwen2-72B
LMArena Text10501203
LMArena Creative Writing10241181
LMArena Multi-Turn9891196

Frequently asked questions

Is Phi 3 Mini 128k Instruct better than Qwen2-72B?

Phi 3 Mini 128k Instruct and Qwen2-72B score almost the same on the Noometry Index (29.7 vs 30.0), so choose on price, context window or the category you care about most.

Is Phi 3 Mini 128k Instruct or Qwen2-72B better for coding?

They score almost the same on coding (28.8 vs 29.1); test both on your own repository before choosing.

How many benchmarks do Phi 3 Mini 128k Instruct and Qwen2-72B share?

19 benchmarks have published results for both models. Phi 3 Mini 128k Instruct has 19 scored results on Noometry and Qwen2-72B has 26.

Related comparisons

Go deeper