Model comparison

Llama 3.2 3B vs Mercury

Mercury is the stronger model overall, scoring 37.6 to 28.9 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 3.2 3B scores higher in 1 category and Mercury in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mercury leads 46.2 to 24.7.
  • Llama 3.2 3B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 3B and Mercury specifications
Llama 3.2 3BMercury
ProviderMetaInception
Noometry Index28.937.6
Released2024-09-24—
WeightsOpenProprietary
Context window131K—
Max output118K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.33—
Results tracked189

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Llama 3.2 3B: 27.6 (#319), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkLlama 3.2 3BMercury
LMArena Coding10981322
BigCodeBench Instruct23.4%—
BigCodeBench Complete28.3%—

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Mercury: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BMercury
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Llama 3.2 3B leads

Llama 3.2 3B: 21.0 (#228), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkLlama 3.2 3BMercury
LMArena Hard Prompts10951285
Kagi LLM Benchmark—21.6%

Math Not comparable

Llama 3.2 3B: 32.4 (#214), Mercury: —

Math benchmarks
BenchmarkLlama 3.2 3BMercury
LMArena Math1126—

Knowledge Not comparable

Llama 3.2 3B: 29.7 (#235), Mercury: —

Knowledge benchmarks
BenchmarkLlama 3.2 3BMercury
LMArena Expert1090—

Multilingual Mercury leads

Llama 3.2 3B: 26.2 (#281), Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkLlama 3.2 3BMercury
LMArena Non-English10191260
LMArena Chinese1017—
LMArena German1056—
LMArena Russian949—

Instruction Following Mercury leads

Llama 3.2 3B: 56.0 (#275), Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BMercury
LMArena Instruction Following10891239

Long Context Mercury leads

Llama 3.2 3B: 33.4 (#261), Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkLlama 3.2 3BMercury
LMArena Longer Query11001266

Writing & Preference Mercury leads

Llama 3.2 3B: 24.7 (#307), Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BMercury
LMArena Text11101282
LMArena Creative Writing10941191
LMArena Multi-Turn11051282
EQ-Bench Creative Writing595—

Frequently asked questions

Is Llama 3.2 3B better than Mercury?

Mercury is the stronger model overall, scoring 37.6 to 28.9 on the Noometry Index.

Is Llama 3.2 3B or Mercury better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 27.6 in the Noometry coding category.

How many benchmarks do Llama 3.2 3B and Mercury share?

8 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper