Model comparison

Mercury 2.5 vs Phi 3 Medium 4k Instruct

Mercury 2.5 and Phi 3 Medium 4k Instruct score almost the same on the Noometry Index (33.5 vs 33.0), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Phi 3 Medium 4k Instruct Microsoft

33.0

Rank #250 Confirmed

Summary

  • The widest gap is in math, where Phi 3 Medium 4k Instruct leads 33.4 to 23.3.
  • Phi 3 Medium 4k Instruct has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and Phi 3 Medium 4k Instruct specifications
Mercury 2.5Phi 3 Medium 4k Instruct
ProviderInceptionMicrosoft
Noometry Index33.533.0
Released2026-09-08—
WeightsProprietaryOpen
Context window260K—
Max output66K—
Input $ / M tokens$0.04—
Output $ / M tokens$0.15—
Results tracked417

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Mercury 2.5: 39.5 (#156), Phi 3 Medium 4k Instruct: 32.9 (#268)

Coding benchmarks
BenchmarkMercury 2.5Phi 3 Medium 4k Instruct
SciCode38.5%—
LMArena Coding—1130
ALE-Bench301.65—

Reasoning Too close to call

Mercury 2.5: 22.4 (#193), Phi 3 Medium 4k Instruct: 21.7 (#214)

Reasoning benchmarks
BenchmarkMercury 2.5Phi 3 Medium 4k Instruct
CritPt0%—
LMArena Hard Prompts—1127

Math Phi 3 Medium 4k Instruct leads

Mercury 2.5: 23.3 (#272), Phi 3 Medium 4k Instruct: 33.4 (#202)

Math benchmarks
BenchmarkMercury 2.5Phi 3 Medium 4k Instruct
ProofBench3%—
LMArena Math—1173

Knowledge Not comparable

Mercury 2.5: —, Phi 3 Medium 4k Instruct: 30.2 (#229)

Knowledge benchmarks
BenchmarkMercury 2.5Phi 3 Medium 4k Instruct
LMArena Expert—1108

Multilingual Not comparable

Mercury 2.5: —, Phi 3 Medium 4k Instruct: 30.6 (#263)

Multilingual benchmarks
BenchmarkMercury 2.5Phi 3 Medium 4k Instruct
LMArena Non-English—1094
LMArena Chinese—1108
LMArena French—1110
LMArena German—1101
LMArena Japanese—1042
LMArena Korean—954
LMArena Russian—1145
LMArena Spanish—1095

Instruction Following Not comparable

Mercury 2.5: —, Phi 3 Medium 4k Instruct: 57.6 (#268)

Instruction Following benchmarks
BenchmarkMercury 2.5Phi 3 Medium 4k Instruct
LMArena Instruction Following—1113

Long Context Not comparable

Mercury 2.5: —, Phi 3 Medium 4k Instruct: 34.0 (#255)

Long Context benchmarks
BenchmarkMercury 2.5Phi 3 Medium 4k Instruct
LMArena Longer Query—1120

Writing & Preference Not comparable

Mercury 2.5: —, Phi 3 Medium 4k Instruct: 34.1 (#272)

Writing & Preference benchmarks
BenchmarkMercury 2.5Phi 3 Medium 4k Instruct
LMArena Text—1138
LMArena Creative Writing—1107
LMArena Multi-Turn—1088

Frequently asked questions

Is Mercury 2.5 better than Phi 3 Medium 4k Instruct?

Mercury 2.5 and Phi 3 Medium 4k Instruct score almost the same on the Noometry Index (33.5 vs 33.0), so choose on price, context window or the category you care about most.

Is Mercury 2.5 or Phi 3 Medium 4k Instruct better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 32.9 in the Noometry coding category.

How many benchmarks do Mercury 2.5 and Phi 3 Medium 4k Instruct share?

0 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and Phi 3 Medium 4k Instruct has 17.

Related comparisons

Go deeper