Model comparison
Claude 2.1 vs Mercury 2.5
Mercury 2.5 is the stronger model overall, scoring 33.5 to 25.2 on the Noometry Index.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in coding, where Mercury 2.5 leads 39.5 to 26.2.
Side by side
| Claude 2.1 | Mercury 2.5 | |
|---|---|---|
| Provider | Anthropic | Inception |
| Noometry Index | 25.2 | 33.5 |
| Released | 2023-11-21 | 2026-09-08 |
| Weights | Proprietary | Proprietary |
| Context window | — | 260K |
| Max output | — | 66K |
| Input $ / M tokens | — | $0.04 |
| Output $ / M tokens | — | $0.15 |
| Results tracked | 7 | 4 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Mercury 2.5 leads
Claude 2.1: 26.2 (#327), Mercury 2.5: 39.5 (#156)
Reasoning Mercury 2.5 leads
Claude 2.1: 21.4 (#221), Mercury 2.5: 22.4 (#193)
| Benchmark | Claude 2.1 | Mercury 2.5 |
|---|---|---|
| CritPt | — | 0% |
| DTBench | 51% | — |
| Epoch Capabilities Index | 119.27 | — |
| ForecastBench | 54.2 | — |
Math Mercury 2.5 leads
Claude 2.1: 10.2 (#315), Mercury 2.5: 23.3 (#272)
| Benchmark | Claude 2.1 | Mercury 2.5 |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 1.9% | — |
| ProofBench | — | 3% |
Knowledge Not comparable
Claude 2.1: 15.4 (#292), Mercury 2.5: —
| Benchmark | Claude 2.1 | Mercury 2.5 |
|---|---|---|
| GPQA Diamond | 33% | — |
| MMLU | 73.5% | — |
Frequently asked questions
Is Claude 2.1 better than Mercury 2.5?
Mercury 2.5 is the stronger model overall, scoring 33.5 to 25.2 on the Noometry Index.
Is Claude 2.1 or Mercury 2.5 better for coding?
Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 26.2 in the Noometry coding category.
How many benchmarks do Claude 2.1 and Mercury 2.5 share?
0 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Mercury 2.5 has 4.