Model comparison

Llama 3.2 90B vs Mercury 2.5

Mercury 2.5 is the stronger model overall, scoring 33.5 to 27.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Mercury 2.5 Inception

33.5

Rank #242 Reported

Summary

  • The widest gap is in math, where Mercury 2.5 leads 23.3 to 11.1.
  • Llama 3.2 90B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 90B and Mercury 2.5 specifications
Llama 3.2 90BMercury 2.5
ProviderMetaInception
Noometry Index27.533.5
Released2024-09-242026-09-08
WeightsOpenProprietary
Context window—260K
Max output—66K
Input $ / M tokens—$0.04
Output $ / M tokens—$0.15
Results tracked94

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Mercury 2.5: 39.5 (#156)

Coding benchmarks
BenchmarkLlama 3.2 90BMercury 2.5
SciCode—38.5%
ALE-Bench—301.65

Agentic & Tool Use Not comparable

Llama 3.2 90B: 30.0 (#80), Mercury 2.5: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BMercury 2.5
BALROG27.3%—

Reasoning Too close to call

Llama 3.2 90B: 21.7 (#217), Mercury 2.5: 22.4 (#193)

Reasoning benchmarks
BenchmarkLlama 3.2 90BMercury 2.5
CritPt—0%
EnigmaEval0.4%—
Epoch Capabilities Index125.5—

Math Mercury 2.5 leads

Llama 3.2 90B: 11.1 (#308), Mercury 2.5: 23.3 (#272)

Math benchmarks
BenchmarkLlama 3.2 90BMercury 2.5
OTIS Mock AIME 2024-20252.6%—
ProofBench—3%
MATH Level 539.4%—

Knowledge Not comparable

Llama 3.2 90B: 21.7 (#274), Mercury 2.5: —

Knowledge benchmarks
BenchmarkLlama 3.2 90BMercury 2.5
GPQA Diamond41%—
MMLU80.3%—

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Mercury 2.5: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BMercury 2.5
LMArena Vision1000—
GeoBench52%—

Frequently asked questions

Is Llama 3.2 90B better than Mercury 2.5?

Mercury 2.5 is the stronger model overall, scoring 33.5 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Mercury 2.5 share?

0 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Mercury 2.5 has 4.

Related comparisons

Go deeper