Model comparison

Llama 3.2 1B vs Mercury 2

Mercury 2 is the stronger model overall, scoring 39.1 to 20.1 on the Noometry Index. Llama 3.2 1B costs 5.3× less per token, which makes it the better buy when Mercury 2's lead doesn't matter for your workload.

Last verified . 11 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Mercury 2 Inception

39.1

Rank #175 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Llama 3.2 1B scores higher in 0 categories and Mercury 2 in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mercury 2 leads 53.8 to 21.3.
  • Llama 3.2 1B is cheaper at $0.027 / $0.20 per million input/output tokens, against $0.25 / $0.75 for Mercury 2.
  • Mercury 2 accepts more context: 128K tokens versus 60K.
  • Llama 3.2 1B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 1B and Mercury 2 specifications
Llama 3.2 1BMercury 2
ProviderMetaInception
Noometry Index20.139.1
Released2024-09-242026-02-20
WeightsOpenProprietary
Context window60K128K
Max output54K50K
Input $ / M tokens$0.027$0.25
Output $ / M tokens$0.20$0.75
Results tracked2217

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2 leads

Llama 3.2 1B: 21.1 (#338), Mercury 2: 33.5 (#255)

Coding benchmarks
BenchmarkLlama 3.2 1BMercury 2
LMArena Coding10701391
LMArena WebDev—1171
SciCode—38.7%
WeirdML—43.2%
BigCodeBench Instruct8.2%—
BigCodeBench Complete11.3%—
ALE-Bench—785.58

Agentic & Tool Use Not comparable

Llama 3.2 1B: 14.6 (#150), Mercury 2: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BMercury 2
Berkeley Function Calling Leaderboard10.8%—
BALROG6.6%—

Reasoning Mercury 2 leads

Llama 3.2 1B: 16.2 (#308), Mercury 2: 23.8 (#170)

Reasoning benchmarks
BenchmarkLlama 3.2 1BMercury 2
LMArena Hard Prompts10441362
CritPt—0.8%
Chess Puzzles0%—
Epoch Capabilities Index101.99—

Math Not comparable

Llama 3.2 1B: 10.4 (#313), Mercury 2: —

Math benchmarks
BenchmarkLlama 3.2 1BMercury 2
OTIS Mock AIME 2024-20250.6%—
LMArena Math1086—

Knowledge Mercury 2 leads

Llama 3.2 1B: 7.2 (#312), Mercury 2: 36.2 (#172)

Knowledge benchmarks
BenchmarkLlama 3.2 1BMercury 2
LMArena Expert10071358
GPQA Diamond23.9%—
Vectara Hallucination Rate—12.3%

Multilingual Mercury 2 leads

Llama 3.2 1B: 23.8 (#292), Mercury 2: 46.6 (#157)

Multilingual benchmarks
BenchmarkLlama 3.2 1BMercury 2
LMArena Non-English9731331
LMArena Chinese9591417
LMArena Russian9411304
LMArena German1014—

Instruction Following Mercury 2 leads

Llama 3.2 1B: 52.4 (#290), Mercury 2: 70.2 (#165)

Instruction Following benchmarks
BenchmarkLlama 3.2 1BMercury 2
LMArena Instruction Following10311329

Long Context Mercury 2 leads

Llama 3.2 1B: 31.9 (#274), Mercury 2: 40.5 (#154)

Long Context benchmarks
BenchmarkLlama 3.2 1BMercury 2
LMArena Longer Query10501330

Writing & Preference Mercury 2 leads

Llama 3.2 1B: 21.3 (#310), Mercury 2: 53.8 (#155)

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BMercury 2
LMArena Text10551355
LMArena Creative Writing10331289
LMArena Multi-Turn10301358
EQ-Bench Creative Writing200—

Frequently asked questions

Is Llama 3.2 1B better than Mercury 2?

Mercury 2 is the stronger model overall, scoring 39.1 to 20.1 on the Noometry Index. Llama 3.2 1B costs 5.3× less per token, which makes it the better buy when Mercury 2's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 1B or Mercury 2?

Llama 3.2 1B is cheaper. It lists at $0.027 per million input tokens and $0.20 per million output tokens; Mercury 2 lists at $0.25 and $0.75.

Is Llama 3.2 1B or Mercury 2 better for coding?

Mercury 2 scores higher on coding benchmarks: 33.5 versus 21.1 in the Noometry coding category.

Which has the bigger context window?

Mercury 2 does, with 128K tokens against 60K.

How many benchmarks do Llama 3.2 1B and Mercury 2 share?

11 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Mercury 2 has 17.

Related comparisons

Go deeper