Model comparison

Mercury vs Qwen3 32B

Qwen3 32B is the stronger model overall, scoring 39.2 to 37.6 on the Noometry Index.

Last verified . 9 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Mercury scores higher in 1 category and Qwen3 32B in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3 32B leads 52.9 to 46.2.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 21.6% for Mercury and 54.9% for Qwen3 32B.
  • Qwen3 32B has downloadable open weights; the other is API-only.

Side by side

Mercury and Qwen3 32B specifications
MercuryQwen3 32B
ProviderInceptionAlibaba (Qwen)
Noometry Index37.639.2
Released—2025-04
WeightsProprietaryOpen
Context window—131K
Max output—16K
Input $ / M tokens—$0.70
Output $ / M tokens—$2.80
Results tracked926

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mercury: 38.7 (#170), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkMercuryQwen3 32B
LMArena Coding13221358
Aider Polyglot—40%
SciCode—35.4%

Agentic & Tool Use Not comparable

Mercury: —, Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkMercuryQwen3 32B
Berkeley Function Calling Leaderboard—48.7%

Reasoning Qwen3 32B leads

Mercury: 17.5 (#293), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkMercuryQwen3 32B
Kagi LLM Benchmark21.6%54.9%
LMArena Hard Prompts12851334
CritPt—0.3%
Chess Puzzles—5%
DTBench—67.5%
LMCA—17.3%
Epoch Capabilities Index—138.51

Math Not comparable

Mercury: —, Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkMercuryQwen3 32B
OTIS Mock AIME 2024-2025—66.9%
LMArena Math—1399

Knowledge Not comparable

Mercury: —, Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkMercuryQwen3 32B
GPQA Diamond—65.7%
Vectara Hallucination Rate—5.9%
LMArena Expert—1362

Multilingual Qwen3 32B leads

Mercury: 41.6 (#206), Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkMercuryQwen3 32B
LMArena Non-English12601317
LMArena Chinese—1357
LMArena German—1341
LMArena Russian—1311

Instruction Following Qwen3 32B leads

Mercury: 65.2 (#224), Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkMercuryQwen3 32B
LMArena Instruction Following12391305

Long Context Qwen3 32B leads

Mercury: 38.4 (#198), Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkMercuryQwen3 32B
LMArena Longer Query12661327
Fiction.LiveBench—74.2%

Writing & Preference Qwen3 32B leads

Mercury: 46.2 (#221), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkMercuryQwen3 32B
LMArena Text12821340
LMArena Creative Writing11911297
LMArena Multi-Turn12821331

Frequently asked questions

Is Mercury better than Qwen3 32B?

Qwen3 32B is the stronger model overall, scoring 39.2 to 37.6 on the Noometry Index.

Is Mercury or Qwen3 32B better for coding?

They score almost the same on coding (38.7 vs 37.7); test both on your own repository before choosing.

How many benchmarks do Mercury and Qwen3 32B share?

9 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper