Model comparison

Mercury vs Qwen3.6 Plus

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 37.6 on the Noometry Index.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Qwen3.6 Plus Alibaba (Qwen)

47.5

Rank #62 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Mercury scores higher in 0 categories and Qwen3.6 Plus in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.6 Plus leads 62.2 to 46.2.

Side by side

Mercury and Qwen3.6 Plus specifications
MercuryQwen3.6 Plus
ProviderInceptionAlibaba (Qwen)
Noometry Index37.647.5
Released—2026-03-31
WeightsProprietaryProprietary
Context window—1M
Max output—66K
Input $ / M tokens—$0.50
Output $ / M tokens—$3
Results tracked937

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Plus leads

Mercury: 38.7 (#170), Qwen3.6 Plus: 40.8 (#130)

Coding benchmarks
BenchmarkMercuryQwen3.6 Plus
LMArena Coding13221467
SWE-bench Verified—57.9%
LMArena WebDev—1461
SciCode—40.7%
ALE-Bench—670.15

Agentic & Tool Use Not comparable

Mercury: —, Qwen3.6 Plus: —

Agentic & Tool Use benchmarks
BenchmarkMercuryQwen3.6 Plus
Vending-Bench 2—5,115

Reasoning Qwen3.6 Plus leads

Mercury: 17.5 (#293), Qwen3.6 Plus: 29.3 (#93)

Reasoning benchmarks
BenchmarkMercuryQwen3.6 Plus
LMArena Hard Prompts12851449
Kagi LLM Benchmark21.6%—
NYT Connections (extended)—60.3%
CritPt—2.9%
Chess Puzzles—17%
Thematic Generalization—59.5%
Mystery Game Puzzles—12%
DTBench—81.9%
LMCA—33.1%
Epoch Capabilities Index—147.65

Math Not comparable

Mercury: —, Qwen3.6 Plus: 51.8 (#54)

Math benchmarks
BenchmarkMercuryQwen3.6 Plus
FrontierMath (Tiers 1-3)—38.2%
OTIS Mock AIME 2024-2025—93.3%
LMArena Math—1450
FrontierMath (Feb 2025 set)—26.2%
FrontierMath Tier 4 (v1)—8.3%

Knowledge Not comparable

Mercury: —, Qwen3.6 Plus: 56.1 (#45)

Knowledge benchmarks
BenchmarkMercuryQwen3.6 Plus
GPQA Diamond—88.4%
SimpleQA Verified—44.1%
LMArena Expert—1454

Multilingual Qwen3.6 Plus leads

Mercury: 41.6 (#206), Qwen3.6 Plus: 53.3 (#70)

Multilingual benchmarks
BenchmarkMercuryQwen3.6 Plus
LMArena Non-English12601424
LMArena Chinese—1477
LMArena French—1455
LMArena German—1452
LMArena Japanese—1389
LMArena Korean—1379
LMArena Russian—1434
LMArena Spanish—1432

Instruction Following Qwen3.6 Plus leads

Mercury: 65.2 (#224), Qwen3.6 Plus: 75.0 (#74)

Instruction Following benchmarks
BenchmarkMercuryQwen3.6 Plus
LMArena Instruction Following12391425

Long Context Qwen3.6 Plus leads

Mercury: 38.4 (#198), Qwen3.6 Plus: 45.2 (#49)

Long Context benchmarks
BenchmarkMercuryQwen3.6 Plus
LMArena Longer Query12661439
CL-bench—20.3%

Writing & Preference Qwen3.6 Plus leads

Mercury: 46.2 (#221), Qwen3.6 Plus: 62.2 (#82)

Writing & Preference benchmarks
BenchmarkMercuryQwen3.6 Plus
LMArena Text12821437
LMArena Creative Writing11911404
LMArena Multi-Turn12821438

Frequently asked questions

Is Mercury better than Qwen3.6 Plus?

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 37.6 on the Noometry Index.

Is Mercury or Qwen3.6 Plus better for coding?

Qwen3.6 Plus scores higher on coding benchmarks: 40.8 versus 38.7 in the Noometry coding category.

How many benchmarks do Mercury and Qwen3.6 Plus share?

8 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Qwen3.6 Plus has 37.

Related comparisons

Go deeper