Model comparison

Mercury vs Qwen2.5 7B Instruct

Mercury is the stronger model overall, scoring 37.6 to 29.0 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • Qwen2.5 7B Instruct has downloadable open weights; the other is API-only.

Side by side

Mercury and Qwen2.5 7B Instruct specifications
MercuryQwen2.5 7B Instruct
ProviderInceptionAlibaba (Qwen)
Noometry Index37.629.0
Released—2024-09
WeightsProprietaryOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.17
Output $ / M tokens—$0.70
Results tracked915

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Mercury: 38.7 (#170), Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
BenchmarkMercuryQwen2.5 7B Instruct
BigCodeBench Instruct—37.6%
LMArena Coding1322—
BigCodeBench Complete—46.1%

Agentic & Tool Use Not comparable

Mercury: —, Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
BenchmarkMercuryQwen2.5 7B Instruct
BALROG—7.8%

Reasoning Mercury leads

Mercury: 17.5 (#293), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
BenchmarkMercuryQwen2.5 7B Instruct
Kagi LLM Benchmark21.6%—
Chess Puzzles—0%
LMArena Hard Prompts1285—
DTBench—47.7%
LMCA—6.4%
Epoch Capabilities Index—118.51

Math Not comparable

Mercury: —, Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
BenchmarkMercuryQwen2.5 7B Instruct
OTIS Mock AIME 2024-2025—2.5%
Omni-MATH—29.4%

Knowledge Not comparable

Mercury: —, Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
BenchmarkMercuryQwen2.5 7B Instruct
GPQA Diamond—35.5%
MMLU-Pro—53.9%
GPQA (HELM)—34.1%
MMLU—72.9%

Multilingual Not comparable

Mercury: 41.6 (#206), Qwen2.5 7B Instruct: —

Multilingual benchmarks
BenchmarkMercuryQwen2.5 7B Instruct
LMArena Non-English1260—

Instruction Following Mercury leads

Mercury: 65.2 (#224), Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
BenchmarkMercuryQwen2.5 7B Instruct
IFEval—74.1%
LMArena Instruction Following1239—

Long Context Not comparable

Mercury: 38.4 (#198), Qwen2.5 7B Instruct: —

Long Context benchmarks
BenchmarkMercuryQwen2.5 7B Instruct
LMArena Longer Query1266—

Writing & Preference Qwen2.5 7B Instruct leads

Mercury: 46.2 (#221), Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
BenchmarkMercuryQwen2.5 7B Instruct
LMArena Text1282—
LMArena Creative Writing1191—
WildBench—73.1%
LMArena Multi-Turn1282—

Frequently asked questions

Is Mercury better than Qwen2.5 7B Instruct?

Mercury is the stronger model overall, scoring 37.6 to 29.0 on the Noometry Index.

Is Mercury or Qwen2.5 7B Instruct better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 36.5 in the Noometry coding category.

How many benchmarks do Mercury and Qwen2.5 7B Instruct share?

0 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper