Model comparison

phi-3-medium 14B vs Qwen3.7 Plus

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 29.7 on the Noometry Index.

Last verified . 2 shared benchmarks.

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Qwen3.7 Plus Alibaba (Qwen)

45.3

Rank #72 Confirmed

Summary

  • They share 2 benchmarks with published results for both. phi-3-medium 14B scores higher in 1 category and Qwen3.7 Plus in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.7 Plus leads 54.9 to 9.1.
  • The biggest single-benchmark swing is GPQA Diamond: 27.6% for phi-3-medium 14B and 87.9% for Qwen3.7 Plus.
  • phi-3-medium 14B has downloadable open weights; the other is API-only.

Side by side

phi-3-medium 14B and Qwen3.7 Plus specifications
phi-3-medium 14BQwen3.7 Plus
ProviderMicrosoftAlibaba (Qwen)
Noometry Index29.745.3
Released2024-04-232026-06-02
WeightsOpenProprietary
Context window—1M
Max output—131K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.60
Results tracked1332

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

phi-3-medium 14B: 36.8 (#201), Qwen3.7 Plus: 36.6 (#206)

Coding benchmarks
Benchmarkphi-3-medium 14BQwen3.7 Plus
FrontierCode—10.2%
SciCode—45.5%
BigCodeBench Instruct37.6%—
LMArena Coding—1473
BigCodeBench Complete48.7%—

Agentic & Tool Use Not comparable

phi-3-medium 14B: —, Qwen3.7 Plus: 21.4 (#138)

Agentic & Tool Use benchmarks
Benchmarkphi-3-medium 14BQwen3.7 Plus
OSWorld 2.0—2.8%

Reasoning Not comparable

phi-3-medium 14B: —, Qwen3.7 Plus: 39.3 (#59)

Reasoning benchmarks
Benchmarkphi-3-medium 14BQwen3.7 Plus
Epoch Capabilities Index121.23147.37
NYT Connections (extended)—74.8%
CritPt—9.1%
Chess Puzzles—24%
LMArena Hard Prompts—1460
Mystery Game Puzzles—17%
DTBench—84%
LMCA—37.6%
Adversarial NLI55.8%—
BIG-Bench Hard81.4%—
HellaSwag82.4%—
WinoGrande81.5%—

Math Qwen3.7 Plus leads

phi-3-medium 14B: 27.3 (#250), Qwen3.7 Plus: 50.5 (#56)

Math benchmarks
Benchmarkphi-3-medium 14BQwen3.7 Plus
FrontierMath (Tiers 1-3)—34.4%
OTIS Mock AIME 2024-2025—93.3%
LMArena Math—1466
MATH Level 517.6%—

Knowledge Qwen3.7 Plus leads

phi-3-medium 14B: 9.1 (#306), Qwen3.7 Plus: 54.9 (#51)

Knowledge benchmarks
Benchmarkphi-3-medium 14BQwen3.7 Plus
GPQA Diamond27.6%87.9%
LMArena Expert—1467
ARC (AI2) Challenge91.6%—
MMLU78%—
OpenBookQA87.4%—
TriviaQA73.9%—

Multimodal Not comparable

phi-3-medium 14B: —, Qwen3.7 Plus: 41.8 (#33)

Multimodal benchmarks
Benchmarkphi-3-medium 14BQwen3.7 Plus
LMArena Vision—1279
LMArena Document—1444

Multilingual Not comparable

phi-3-medium 14B: —, Qwen3.7 Plus: 54.8 (#38)

Multilingual benchmarks
Benchmarkphi-3-medium 14BQwen3.7 Plus
LMArena Non-English—1445
LMArena Chinese—1510
LMArena French—1473
LMArena German—1471
LMArena Japanese—1413
LMArena Korean—1415
LMArena Russian—1457
LMArena Spanish—1457

Instruction Following Not comparable

phi-3-medium 14B: —, Qwen3.7 Plus: 75.8 (#52)

Instruction Following benchmarks
Benchmarkphi-3-medium 14BQwen3.7 Plus
LMArena Instruction Following—1440

Long Context Not comparable

phi-3-medium 14B: —, Qwen3.7 Plus: 44.5 (#65)

Long Context benchmarks
Benchmarkphi-3-medium 14BQwen3.7 Plus
LMArena Longer Query—1455

Writing & Preference Not comparable

phi-3-medium 14B: —, Qwen3.7 Plus: 64.3 (#56)

Writing & Preference benchmarks
Benchmarkphi-3-medium 14BQwen3.7 Plus
LMArena Text—1455
LMArena Creative Writing—1439
LMArena Multi-Turn—1460

Frequently asked questions

Is phi-3-medium 14B better than Qwen3.7 Plus?

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 29.7 on the Noometry Index.

Is phi-3-medium 14B or Qwen3.7 Plus better for coding?

They score almost the same on coding (36.8 vs 36.6); test both on your own repository before choosing.

How many benchmarks do phi-3-medium 14B and Qwen3.7 Plus share?

2 benchmarks have published results for both models. phi-3-medium 14B has 13 scored results on Noometry and Qwen3.7 Plus has 32.

Related comparisons

Go deeper