Model comparison

Mercury vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 37.6 on the Noometry Index.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Mercury scores higher in 0 categories and Qwen3.5 Max Preview in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 Max Preview leads 66.0 to 46.2.

Side by side

Mercury and Qwen3.5 Max Preview specifications
MercuryQwen3.5 Max Preview
ProviderInceptionAlibaba (Qwen)
Noometry Index37.645.3
Released——
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked917

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Mercury: 38.7 (#170), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkMercuryQwen3.5 Max Preview
LMArena Coding13221487

Reasoning Qwen3.5 Max Preview leads

Mercury: 17.5 (#293), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkMercuryQwen3.5 Max Preview
LMArena Hard Prompts12851483
Kagi LLM Benchmark21.6%—

Math Not comparable

Mercury: —, Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkMercuryQwen3.5 Max Preview
LMArena Math—1474

Knowledge Not comparable

Mercury: —, Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkMercuryQwen3.5 Max Preview
LMArena Expert—1489

Multilingual Qwen3.5 Max Preview leads

Mercury: 41.6 (#206), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkMercuryQwen3.5 Max Preview
LMArena Non-English12601465
LMArena Chinese—1534
LMArena French—1484
LMArena German—1487
LMArena Japanese—1495
LMArena Korean—1438
LMArena Russian—1471
LMArena Spanish—1470

Instruction Following Qwen3.5 Max Preview leads

Mercury: 65.2 (#224), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkMercuryQwen3.5 Max Preview
LMArena Instruction Following12391467

Long Context Qwen3.5 Max Preview leads

Mercury: 38.4 (#198), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkMercuryQwen3.5 Max Preview
LMArena Longer Query12661476

Writing & Preference Qwen3.5 Max Preview leads

Mercury: 46.2 (#221), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkMercuryQwen3.5 Max Preview
LMArena Text12821470
LMArena Creative Writing11911464
LMArena Multi-Turn12821478

Frequently asked questions

Is Mercury better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 37.6 on the Noometry Index.

Is Mercury or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 38.7 in the Noometry coding category.

How many benchmarks do Mercury and Qwen3.5 Max Preview share?

8 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper