Model comparison

Llama 3.2 90B vs Mercury

Mercury is the stronger model overall, scoring 37.6 to 27.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • The widest gap is in reasoning, where Llama 3.2 90B leads 21.7 to 17.5.
  • Llama 3.2 90B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 90B and Mercury specifications
Llama 3.2 90BMercury
ProviderMetaInception
Noometry Index27.537.6
Released2024-09-24—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked99

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkLlama 3.2 90BMercury
LMArena Coding—1322

Agentic & Tool Use Not comparable

Llama 3.2 90B: 30.0 (#80), Mercury: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BMercury
BALROG27.3%—

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkLlama 3.2 90BMercury
Kagi LLM Benchmark—21.6%
EnigmaEval0.4%—
LMArena Hard Prompts—1285
Epoch Capabilities Index125.5—

Math Not comparable

Llama 3.2 90B: 11.1 (#308), Mercury: —

Math benchmarks
BenchmarkLlama 3.2 90BMercury
OTIS Mock AIME 2024-20252.6%—
MATH Level 539.4%—

Knowledge Not comparable

Llama 3.2 90B: 21.7 (#274), Mercury: —

Knowledge benchmarks
BenchmarkLlama 3.2 90BMercury
GPQA Diamond41%—
MMLU80.3%—

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Mercury: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BMercury
LMArena Vision1000—
GeoBench52%—

Multilingual Not comparable

Llama 3.2 90B: —, Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkLlama 3.2 90BMercury
LMArena Non-English—1260

Instruction Following Not comparable

Llama 3.2 90B: —, Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BMercury
LMArena Instruction Following—1239

Long Context Not comparable

Llama 3.2 90B: —, Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkLlama 3.2 90BMercury
LMArena Longer Query—1266

Writing & Preference Not comparable

Llama 3.2 90B: —, Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BMercury
LMArena Text—1282
LMArena Creative Writing—1191
LMArena Multi-Turn—1282

Frequently asked questions

Is Llama 3.2 90B better than Mercury?

Mercury is the stronger model overall, scoring 37.6 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Mercury share?

0 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper