Model comparison

Claude 2.1 vs Mercury 2.5

Mercury 2.5 is the stronger model overall, scoring 33.5 to 25.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Mercury 2.5 Inception

33.5

Rank #242 Reported

Summary

  • The widest gap is in coding, where Mercury 2.5 leads 39.5 to 26.2.

Side by side

Claude 2.1 and Mercury 2.5 specifications
Claude 2.1Mercury 2.5
ProviderAnthropicInception
Noometry Index25.233.5
Released2023-11-212026-09-08
WeightsProprietaryProprietary
Context window—260K
Max output—66K
Input $ / M tokens—$0.04
Output $ / M tokens—$0.15
Results tracked74

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Claude 2.1: 26.2 (#327), Mercury 2.5: 39.5 (#156)

Coding benchmarks
BenchmarkClaude 2.1Mercury 2.5
SciCode—38.5%
WeirdML7.1%—
ALE-Bench—301.65

Reasoning Mercury 2.5 leads

Claude 2.1: 21.4 (#221), Mercury 2.5: 22.4 (#193)

Reasoning benchmarks
BenchmarkClaude 2.1Mercury 2.5
CritPt—0%
DTBench51%—
Epoch Capabilities Index119.27—
ForecastBench54.2—

Math Mercury 2.5 leads

Claude 2.1: 10.2 (#315), Mercury 2.5: 23.3 (#272)

Math benchmarks
BenchmarkClaude 2.1Mercury 2.5
OTIS Mock AIME 2024-20251.9%—
ProofBench—3%

Knowledge Not comparable

Claude 2.1: 15.4 (#292), Mercury 2.5: —

Knowledge benchmarks
BenchmarkClaude 2.1Mercury 2.5
GPQA Diamond33%—
MMLU73.5%—

Frequently asked questions

Is Claude 2.1 better than Mercury 2.5?

Mercury 2.5 is the stronger model overall, scoring 33.5 to 25.2 on the Noometry Index.

Is Claude 2.1 or Mercury 2.5 better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Mercury 2.5 share?

0 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Mercury 2.5 has 4.

Related comparisons

Go deeper