Model comparison

Mercury 2 vs Qwen1.5-14B

Mercury 2 is the stronger model overall, scoring 39.1 to 32.7 on the Noometry Index.

Last verified . 11 shared benchmarks.

Mercury 2 Inception

39.1

Rank #175 Confirmed

Qwen1.5-14B Alibaba (Qwen)

32.7

Rank #253 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Mercury 2 scores higher in 7 categories and Qwen1.5-14B in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mercury 2 leads 53.8 to 33.6.
  • Qwen1.5-14B has downloadable open weights; the other is API-only.

Side by side

Mercury 2 and Qwen1.5-14B specifications
Mercury 2Qwen1.5-14B
ProviderInceptionAlibaba (Qwen)
Noometry Index39.132.7
Released2026-02-202024-02-04
WeightsProprietaryOpen
Context window128K—
Max output50K—
Input $ / M tokens$0.25—
Output $ / M tokens$0.75—
Results tracked1717

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mercury 2: 33.5 (#255), Qwen1.5-14B: 33.1 (#263)

Coding benchmarks
BenchmarkMercury 2Qwen1.5-14B
LMArena Coding13911138
LMArena WebDev1171—
SciCode38.7%—
WeirdML43.2%—
ALE-Bench785.58—

Reasoning Mercury 2 leads

Mercury 2: 23.8 (#170), Qwen1.5-14B: 21.4 (#223)

Reasoning benchmarks
BenchmarkMercury 2Qwen1.5-14B
LMArena Hard Prompts13621113
CritPt0.8%—

Math Not comparable

Mercury 2: —, Qwen1.5-14B: 32.4 (#215)

Math benchmarks
BenchmarkMercury 2Qwen1.5-14B
LMArena Math—1125

Knowledge Mercury 2 leads

Mercury 2: 36.2 (#172), Qwen1.5-14B: 29.8 (#232)

Knowledge benchmarks
BenchmarkMercury 2Qwen1.5-14B
LMArena Expert13581094
Vectara Hallucination Rate12.3%—
MMLU—68.6%

Multilingual Mercury 2 leads

Mercury 2: 46.6 (#157), Qwen1.5-14B: 30.7 (#262)

Multilingual benchmarks
BenchmarkMercury 2Qwen1.5-14B
LMArena Non-English13311095
LMArena Chinese14171147
LMArena Russian13041046
LMArena French—1116
LMArena German—1043
LMArena Japanese—1019
LMArena Spanish—1085

Instruction Following Mercury 2 leads

Mercury 2: 70.2 (#165), Qwen1.5-14B: 56.8 (#271)

Instruction Following benchmarks
BenchmarkMercury 2Qwen1.5-14B
LMArena Instruction Following13291102

Long Context Mercury 2 leads

Mercury 2: 40.5 (#154), Qwen1.5-14B: 33.7 (#257)

Long Context benchmarks
BenchmarkMercury 2Qwen1.5-14B
LMArena Longer Query13301113

Writing & Preference Mercury 2 leads

Mercury 2: 53.8 (#155), Qwen1.5-14B: 33.6 (#276)

Writing & Preference benchmarks
BenchmarkMercury 2Qwen1.5-14B
LMArena Text13551128
LMArena Creative Writing12891091
LMArena Multi-Turn13581110

Frequently asked questions

Is Mercury 2 better than Qwen1.5-14B?

Mercury 2 is the stronger model overall, scoring 39.1 to 32.7 on the Noometry Index.

Is Mercury 2 or Qwen1.5-14B better for coding?

They score almost the same on coding (33.5 vs 33.1); test both on your own repository before choosing.

How many benchmarks do Mercury 2 and Qwen1.5-14B share?

11 benchmarks have published results for both models. Mercury 2 has 17 scored results on Noometry and Qwen1.5-14B has 17.

Related comparisons

Go deeper