Model comparison

Mercury vs Qwen2.5 72B Instruct

Mercury is the stronger model overall, scoring 37.6 to 31.9 on the Noometry Index.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Mercury scores higher in 2 categories and Qwen2.5 72B Instruct in 4 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Mercury leads 38.7 to 33.2.
  • Qwen2.5 72B Instruct has downloadable open weights; the other is API-only.

Side by side

Mercury and Qwen2.5 72B Instruct specifications
MercuryQwen2.5 72B Instruct
ProviderInceptionAlibaba (Qwen)
Noometry Index37.631.9
Released—2024-09
WeightsProprietaryOpen
Context window—131K
Max output—8K
Input $ / M tokens—$1.40
Output $ / M tokens—$5.60
Results tracked943

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Mercury: 38.7 (#170), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkMercuryQwen2.5 72B Instruct
LMArena Coding13221292
WeirdML—16%
BigCodeBench Instruct—45.8%
BigCodeBench Complete—55.9%

Agentic & Tool Use Not comparable

Mercury: —, Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkMercuryQwen2.5 72B Instruct
TheAgentCompany—5.7%
BALROG—16.2%
METR Time Horizons—35.8%

Reasoning Qwen2.5 72B Instruct leads

Mercury: 17.5 (#293), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkMercuryQwen2.5 72B Instruct
LMArena Hard Prompts12851271
Kagi LLM Benchmark21.6%—
DTBench—62.9%
LMCA—13.4%
BIG-Bench Hard—79.8%
Epoch Capabilities Index—129
ForecastBench—57.5
HellaSwag—84.8%
PIQA—82.6%
WinoGrande—82.3%

Math Not comparable

Mercury: —, Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkMercuryQwen2.5 72B Instruct
OTIS Mock AIME 2024-2025—8.1%
Omni-MATH—33%
LMArena Math—1283
MATH Level 5—63.2%

Knowledge Not comparable

Mercury: —, Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkMercuryQwen2.5 72B Instruct
GPQA Diamond—49.1%
MMLU-Pro—63.1%
Confabulations—19.1%
GPQA (HELM)—42.6%
LMArena Expert—1245
ARC (AI2) Challenge—94.5%
MMLU—85.3%
TriviaQA—71.9%

Multilingual Too close to call

Mercury: 41.6 (#206), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkMercuryQwen2.5 72B Instruct
LMArena Non-English12601252
LMArena Chinese—1272
LMArena French—1280
LMArena German—1234
LMArena Japanese—1180
LMArena Korean—1188
LMArena Russian—1264
LMArena Spanish—1256

Instruction Following Too close to call

Mercury: 65.2 (#224), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkMercuryQwen2.5 72B Instruct
LMArena Instruction Following12391254
IFEval—80.6%

Long Context Too close to call

Mercury: 38.4 (#198), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkMercuryQwen2.5 72B Instruct
LMArena Longer Query12661282

Writing & Preference Too close to call

Mercury: 46.2 (#221), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkMercuryQwen2.5 72B Instruct
LMArena Text12821269
LMArena Creative Writing11911221
LMArena Multi-Turn12821272
WildBench—80.2%

Frequently asked questions

Is Mercury better than Qwen2.5 72B Instruct?

Mercury is the stronger model overall, scoring 37.6 to 31.9 on the Noometry Index.

Is Mercury or Qwen2.5 72B Instruct better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 33.2 in the Noometry coding category.

How many benchmarks do Mercury and Qwen2.5 72B Instruct share?

8 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper