Model comparison

Mercury vs Qwen2.5 32B Instruct

Mercury is the stronger model overall, scoring 37.6 to 30.1 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Qwen2.5 32B Instruct Alibaba (Qwen)

30.1

Rank #297 Confirmed

Summary

  • Qwen2.5 32B Instruct has downloadable open weights; the other is API-only.

Side by side

Mercury and Qwen2.5 32B Instruct specifications
MercuryQwen2.5 32B Instruct
ProviderInceptionAlibaba (Qwen)
Noometry Index37.630.1
Released—2024-09
WeightsProprietaryOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.70
Output $ / M tokens—$2.80
Results tracked97

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mercury: 38.7 (#170), Qwen2.5 32B Instruct: 38.7 (#169)

Coding benchmarks
BenchmarkMercuryQwen2.5 32B Instruct
BigCodeBench Instruct—45%
LMArena Coding1322—
BigCodeBench Complete—52.3%

Reasoning Qwen2.5 32B Instruct leads

Mercury: 17.5 (#293), Qwen2.5 32B Instruct: 19.2 (#266)

Reasoning benchmarks
BenchmarkMercuryQwen2.5 32B Instruct
Kagi LLM Benchmark21.6%—
Chess Puzzles—0%
LMArena Hard Prompts1285—
Epoch Capabilities Index—128.52

Math Not comparable

Mercury: —, Qwen2.5 32B Instruct: 16.2 (#296)

Math benchmarks
BenchmarkMercuryQwen2.5 32B Instruct
OTIS Mock AIME 2024-2025—7.4%
MATH Level 5—56.1%

Knowledge Not comparable

Mercury: —, Qwen2.5 32B Instruct: 24.9 (#266)

Knowledge benchmarks
BenchmarkMercuryQwen2.5 32B Instruct
GPQA Diamond—46.1%

Multilingual Not comparable

Mercury: 41.6 (#206), Qwen2.5 32B Instruct: —

Multilingual benchmarks
BenchmarkMercuryQwen2.5 32B Instruct
LMArena Non-English1260—

Instruction Following Not comparable

Mercury: 65.2 (#224), Qwen2.5 32B Instruct: —

Instruction Following benchmarks
BenchmarkMercuryQwen2.5 32B Instruct
LMArena Instruction Following1239—

Long Context Not comparable

Mercury: 38.4 (#198), Qwen2.5 32B Instruct: —

Long Context benchmarks
BenchmarkMercuryQwen2.5 32B Instruct
LMArena Longer Query1266—

Writing & Preference Not comparable

Mercury: 46.2 (#221), Qwen2.5 32B Instruct: —

Writing & Preference benchmarks
BenchmarkMercuryQwen2.5 32B Instruct
LMArena Text1282—
LMArena Creative Writing1191—
LMArena Multi-Turn1282—

Frequently asked questions

Is Mercury better than Qwen2.5 32B Instruct?

Mercury is the stronger model overall, scoring 37.6 to 30.1 on the Noometry Index.

Is Mercury or Qwen2.5 32B Instruct better for coding?

They score almost the same on coding (38.7 vs 38.7); test both on your own repository before choosing.

How many benchmarks do Mercury and Qwen2.5 32B Instruct share?

0 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Qwen2.5 32B Instruct has 7.

Related comparisons

Go deeper