Model comparison

Mercury 2 vs Qwen Plus

Mercury 2 is the stronger model overall, scoring 39.1 to 37.1 on the Noometry Index.

Last verified . 11 shared benchmarks.

Mercury 2 Inception

39.1

Rank #175 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Mercury 2 scores higher in 5 categories and Qwen Plus in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mercury 2 leads 36.2 to 27.4.
  • Mercury 2 is cheaper at $0.25 / $0.75 per million input/output tokens, against $0.40 / $1.20 for Qwen Plus.
  • Qwen Plus accepts more context: 1M tokens versus 128K.

Side by side

Mercury 2 and Qwen Plus specifications
Mercury 2Qwen Plus
ProviderInceptionAlibaba (Qwen)
Noometry Index39.137.1
Released2026-02-202024-01-25
WeightsProprietaryProprietary
Context window128K1M
Max output50K33K
Input $ / M tokens$0.25$0.40
Output $ / M tokens$0.75$1.20
Results tracked1720

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Plus leads

Mercury 2: 33.5 (#255), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkMercury 2Qwen Plus
LMArena Coding13911328
LMArena WebDev1171—
SciCode38.7%—
WeirdML43.2%—
ALE-Bench785.58—

Reasoning Qwen Plus leads

Mercury 2: 23.8 (#170), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkMercury 2Qwen Plus
LMArena Hard Prompts13621317
Kagi LLM Benchmark—63.3%
CritPt0.8%—
DTBench—81.1%
LMCA—24%

Math Not comparable

Mercury 2: —, Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkMercury 2Qwen Plus
OTIS Mock AIME 2024-2025—17.8%
LMArena Math—1326
MATH Level 5—65.3%
FrontierMath (Feb 2025 set)—1.7%

Knowledge Mercury 2 leads

Mercury 2: 36.2 (#172), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkMercury 2Qwen Plus
LMArena Expert13581328
GPQA Diamond—48.1%
Vectara Hallucination Rate12.3%—

Multilingual Mercury 2 leads

Mercury 2: 46.6 (#157), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkMercury 2Qwen Plus
LMArena Non-English13311310
LMArena Chinese14171347
LMArena Russian13041323
LMArena Japanese—1251

Instruction Following Mercury 2 leads

Mercury 2: 70.2 (#165), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkMercury 2Qwen Plus
LMArena Instruction Following13291303

Long Context Too close to call

Mercury 2: 40.5 (#154), Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkMercury 2Qwen Plus
LMArena Longer Query13301324

Writing & Preference Mercury 2 leads

Mercury 2: 53.8 (#155), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkMercury 2Qwen Plus
LMArena Text13551326
LMArena Creative Writing12891293
LMArena Multi-Turn13581336

Frequently asked questions

Is Mercury 2 better than Qwen Plus?

Mercury 2 is the stronger model overall, scoring 39.1 to 37.1 on the Noometry Index.

Which is cheaper, Mercury 2 or Qwen Plus?

Mercury 2 is cheaper. It lists at $0.25 per million input tokens and $0.75 per million output tokens; Qwen Plus lists at $0.40 and $1.20.

Is Mercury 2 or Qwen Plus better for coding?

Qwen Plus scores higher on coding benchmarks: 38.9 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

Qwen Plus does, with 1M tokens against 128K.

How many benchmarks do Mercury 2 and Qwen Plus share?

11 benchmarks have published results for both models. Mercury 2 has 17 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper