Model comparison

Llama 3.2 3B vs Mercury 2

Mercury 2 is the stronger model overall, scoring 39.1 to 28.9 on the Noometry Index. Llama 3.2 3B costs 3.1× less per token, which makes it the better buy when Mercury 2's lead doesn't matter for your workload.

Last verified . 11 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Mercury 2 Inception

39.1

Rank #175 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Llama 3.2 3B scores higher in 0 categories and Mercury 2 in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mercury 2 leads 53.8 to 24.7.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $0.25 / $0.75 for Mercury 2.
  • Llama 3.2 3B accepts more context: 131K tokens versus 128K.
  • Llama 3.2 3B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 3B and Mercury 2 specifications
Llama 3.2 3BMercury 2
ProviderMetaInception
Noometry Index28.939.1
Released2024-09-242026-02-20
WeightsOpenProprietary
Context window131K128K
Max output118K50K
Input $ / M tokens$0.05$0.25
Output $ / M tokens$0.33$0.75
Results tracked1817

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2 leads

Llama 3.2 3B: 27.6 (#319), Mercury 2: 33.5 (#255)

Coding benchmarks
BenchmarkLlama 3.2 3BMercury 2
LMArena Coding10981391
LMArena WebDev—1171
SciCode—38.7%
WeirdML—43.2%
BigCodeBench Instruct23.4%—
BigCodeBench Complete28.3%—
ALE-Bench—785.58

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Mercury 2: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BMercury 2
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Mercury 2 leads

Llama 3.2 3B: 21.0 (#228), Mercury 2: 23.8 (#170)

Reasoning benchmarks
BenchmarkLlama 3.2 3BMercury 2
LMArena Hard Prompts10951362
CritPt—0.8%

Math Not comparable

Llama 3.2 3B: 32.4 (#214), Mercury 2: —

Math benchmarks
BenchmarkLlama 3.2 3BMercury 2
LMArena Math1126—

Knowledge Mercury 2 leads

Llama 3.2 3B: 29.7 (#235), Mercury 2: 36.2 (#172)

Knowledge benchmarks
BenchmarkLlama 3.2 3BMercury 2
LMArena Expert10901358
Vectara Hallucination Rate—12.3%

Multilingual Mercury 2 leads

Llama 3.2 3B: 26.2 (#281), Mercury 2: 46.6 (#157)

Multilingual benchmarks
BenchmarkLlama 3.2 3BMercury 2
LMArena Non-English10191331
LMArena Chinese10171417
LMArena Russian9491304
LMArena German1056—

Instruction Following Mercury 2 leads

Llama 3.2 3B: 56.0 (#275), Mercury 2: 70.2 (#165)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BMercury 2
LMArena Instruction Following10891329

Long Context Mercury 2 leads

Llama 3.2 3B: 33.4 (#261), Mercury 2: 40.5 (#154)

Long Context benchmarks
BenchmarkLlama 3.2 3BMercury 2
LMArena Longer Query11001330

Writing & Preference Mercury 2 leads

Llama 3.2 3B: 24.7 (#307), Mercury 2: 53.8 (#155)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BMercury 2
LMArena Text11101355
LMArena Creative Writing10941289
LMArena Multi-Turn11051358
EQ-Bench Creative Writing595—

Frequently asked questions

Is Llama 3.2 3B better than Mercury 2?

Mercury 2 is the stronger model overall, scoring 39.1 to 28.9 on the Noometry Index. Llama 3.2 3B costs 3.1× less per token, which makes it the better buy when Mercury 2's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 3B or Mercury 2?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Mercury 2 lists at $0.25 and $0.75.

Is Llama 3.2 3B or Mercury 2 better for coding?

Mercury 2 scores higher on coding benchmarks: 33.5 versus 27.6 in the Noometry coding category.

Which has the bigger context window?

Llama 3.2 3B does, with 131K tokens against 128K.

How many benchmarks do Llama 3.2 3B and Mercury 2 share?

11 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Mercury 2 has 17.

Related comparisons

Go deeper