Model comparison

Mercury 2.5 vs Qwen3.5 27B

Qwen3.5 27B is the stronger model overall, scoring 41.9 to 33.5 on the Noometry Index. Mercury 2.5 costs 12× less per token, which makes it the better buy when Qwen3.5 27B's lead doesn't matter for your workload.

Last verified . 1 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Qwen3.5 27B Alibaba (Qwen)

41.9

Rank #127 Confirmed

Summary

  • They share 1 benchmark with published results for both. Mercury 2.5 scores higher in 1 category and Qwen3.5 27B in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.5 27B leads 38.8 to 23.3.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.30 / $2.40 for Qwen3.5 27B.
  • Qwen3.5 27B accepts more context: 262K tokens versus 260K.
  • Qwen3.5 27B has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and Qwen3.5 27B specifications
Mercury 2.5Qwen3.5 27B
ProviderInceptionAlibaba (Qwen)
Noometry Index33.541.9
Released2026-09-082026-02-23
WeightsProprietaryOpen
Context window260K262K
Max output66K66K
Input $ / M tokens$0.04$0.30
Output $ / M tokens$0.15$2.40
Results tracked428

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mercury 2.5: 39.5 (#156), Qwen3.5 27B: 38.9 (#168)

Coding benchmarks
BenchmarkMercury 2.5Qwen3.5 27B
ALE-Bench301.65349.45
LMArena WebDev—1358
SciCode38.5%—
WeirdML—39.5%
LMArena Coding—1427

Agentic & Tool Use Not comparable

Mercury 2.5: —, Qwen3.5 27B: —

Agentic & Tool Use benchmarks
BenchmarkMercury 2.5Qwen3.5 27B
Vending-Bench 2—201.98

Reasoning Qwen3.5 27B leads

Mercury 2.5: 22.4 (#193), Qwen3.5 27B: 27.5 (#117)

Reasoning benchmarks
BenchmarkMercury 2.5Qwen3.5 27B
NYT Connections (extended)—47.9%
CritPt0%—
Thematic Generalization—45.5%
LMArena Hard Prompts—1414
DTBench—82.4%
LMCA—34%

Math Qwen3.5 27B leads

Mercury 2.5: 23.3 (#272), Qwen3.5 27B: 38.8 (#127)

Math benchmarks
BenchmarkMercury 2.5Qwen3.5 27B
MathArena Final-Answer Competitions—56.7%
ProofBench3%—
LMArena Math—1429

Knowledge Not comparable

Mercury 2.5: —, Qwen3.5 27B: 38.0 (#150)

Knowledge benchmarks
BenchmarkMercury 2.5Qwen3.5 27B
Vectara Hallucination Rate—12.1%
LMArena Expert—1428

Multimodal Not comparable

Mercury 2.5: —, Qwen3.5 27B: 39.4 (#59)

Multimodal benchmarks
BenchmarkMercury 2.5Qwen3.5 27B
LMArena Vision—1241

Multilingual Not comparable

Mercury 2.5: —, Qwen3.5 27B: 50.8 (#115)

Multilingual benchmarks
BenchmarkMercury 2.5Qwen3.5 27B
LMArena Non-English—1390
LMArena Chinese—1478
LMArena French—1410
LMArena German—1393
LMArena Japanese—1345
LMArena Korean—1358
LMArena Russian—1390
LMArena Spanish—1407

Instruction Following Not comparable

Mercury 2.5: —, Qwen3.5 27B: 73.5 (#119)

Instruction Following benchmarks
BenchmarkMercury 2.5Qwen3.5 27B
LMArena Instruction Following—1393

Long Context Not comparable

Mercury 2.5: —, Qwen3.5 27B: 43.1 (#106)

Long Context benchmarks
BenchmarkMercury 2.5Qwen3.5 27B
LMArena Longer Query—1413

Writing & Preference Not comparable

Mercury 2.5: —, Qwen3.5 27B: 59.3 (#111)

Writing & Preference benchmarks
BenchmarkMercury 2.5Qwen3.5 27B
LMArena Text—1409
LMArena Creative Writing—1362
LMArena Multi-Turn—1410

Frequently asked questions

Is Mercury 2.5 better than Qwen3.5 27B?

Qwen3.5 27B is the stronger model overall, scoring 41.9 to 33.5 on the Noometry Index. Mercury 2.5 costs 12× less per token, which makes it the better buy when Qwen3.5 27B's lead doesn't matter for your workload.

Which is cheaper, Mercury 2.5 or Qwen3.5 27B?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; Qwen3.5 27B lists at $0.30 and $2.40.

Is Mercury 2.5 or Qwen3.5 27B better for coding?

They score almost the same on coding (39.5 vs 38.9); test both on your own repository before choosing.

Which has the bigger context window?

Qwen3.5 27B does, with 262K tokens against 260K.

How many benchmarks do Mercury 2.5 and Qwen3.5 27B share?

1 benchmark has published results for both models. Mercury 2.5 has 4 scored results on Noometry and Qwen3.5 27B has 28.

Related comparisons

Go deeper