Model comparison
Llama 3.2 1B vs Mercury 2
Mercury 2 is the stronger model overall, scoring 39.1 to 20.1 on the Noometry Index. Llama 3.2 1B costs 5.3× less per token, which makes it the better buy when Mercury 2's lead doesn't matter for your workload.
Last verified . 11 shared benchmarks.
Summary
- They share 11 benchmarks with published results for both. Llama 3.2 1B scores higher in 0 categories and Mercury 2 in 7 categories; 7 gaps are clear of the uncertainty.
- The widest gap is in writing & preference, where Mercury 2 leads 53.8 to 21.3.
- Llama 3.2 1B is cheaper at $0.027 / $0.20 per million input/output tokens, against $0.25 / $0.75 for Mercury 2.
- Mercury 2 accepts more context: 128K tokens versus 60K.
- Llama 3.2 1B has downloadable open weights; the other is API-only.
Side by side
| Llama 3.2 1B | Mercury 2 | |
|---|---|---|
| Provider | Meta | Inception |
| Noometry Index | 20.1 | 39.1 |
| Released | 2024-09-24 | 2026-02-20 |
| Weights | Open | Proprietary |
| Context window | 60K | 128K |
| Max output | 54K | 50K |
| Input $ / M tokens | $0.027 | $0.25 |
| Output $ / M tokens | $0.20 | $0.75 |
| Results tracked | 22 | 17 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Mercury 2 leads
Llama 3.2 1B: 21.1 (#338), Mercury 2: 33.5 (#255)
| Benchmark | Llama 3.2 1B | Mercury 2 |
|---|---|---|
| LMArena Coding | 1070 | 1391 |
| LMArena WebDev | — | 1171 |
| SciCode | — | 38.7% |
| WeirdML | — | 43.2% |
| BigCodeBench Instruct | 8.2% | — |
| BigCodeBench Complete | 11.3% | — |
| ALE-Bench | — | 785.58 |
Agentic & Tool Use Not comparable
Llama 3.2 1B: 14.6 (#150), Mercury 2: —
| Benchmark | Llama 3.2 1B | Mercury 2 |
|---|---|---|
| Berkeley Function Calling Leaderboard | 10.8% | — |
| BALROG | 6.6% | — |
Reasoning Mercury 2 leads
Llama 3.2 1B: 16.2 (#308), Mercury 2: 23.8 (#170)
| Benchmark | Llama 3.2 1B | Mercury 2 |
|---|---|---|
| LMArena Hard Prompts | 1044 | 1362 |
| CritPt | — | 0.8% |
| Chess Puzzles | 0% | — |
| Epoch Capabilities Index | 101.99 | — |
Math Not comparable
Llama 3.2 1B: 10.4 (#313), Mercury 2: —
| Benchmark | Llama 3.2 1B | Mercury 2 |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 0.6% | — |
| LMArena Math | 1086 | — |
Knowledge Mercury 2 leads
Llama 3.2 1B: 7.2 (#312), Mercury 2: 36.2 (#172)
| Benchmark | Llama 3.2 1B | Mercury 2 |
|---|---|---|
| LMArena Expert | 1007 | 1358 |
| GPQA Diamond | 23.9% | — |
| Vectara Hallucination Rate | — | 12.3% |
Multilingual Mercury 2 leads
Llama 3.2 1B: 23.8 (#292), Mercury 2: 46.6 (#157)
| Benchmark | Llama 3.2 1B | Mercury 2 |
|---|---|---|
| LMArena Non-English | 973 | 1331 |
| LMArena Chinese | 959 | 1417 |
| LMArena Russian | 941 | 1304 |
| LMArena German | 1014 | — |
Instruction Following Mercury 2 leads
Llama 3.2 1B: 52.4 (#290), Mercury 2: 70.2 (#165)
| Benchmark | Llama 3.2 1B | Mercury 2 |
|---|---|---|
| LMArena Instruction Following | 1031 | 1329 |
Long Context Mercury 2 leads
Llama 3.2 1B: 31.9 (#274), Mercury 2: 40.5 (#154)
| Benchmark | Llama 3.2 1B | Mercury 2 |
|---|---|---|
| LMArena Longer Query | 1050 | 1330 |
Writing & Preference Mercury 2 leads
Llama 3.2 1B: 21.3 (#310), Mercury 2: 53.8 (#155)
| Benchmark | Llama 3.2 1B | Mercury 2 |
|---|---|---|
| LMArena Text | 1055 | 1355 |
| LMArena Creative Writing | 1033 | 1289 |
| LMArena Multi-Turn | 1030 | 1358 |
| EQ-Bench Creative Writing | 200 | — |
Frequently asked questions
Is Llama 3.2 1B better than Mercury 2?
Mercury 2 is the stronger model overall, scoring 39.1 to 20.1 on the Noometry Index. Llama 3.2 1B costs 5.3× less per token, which makes it the better buy when Mercury 2's lead doesn't matter for your workload.
Which is cheaper, Llama 3.2 1B or Mercury 2?
Llama 3.2 1B is cheaper. It lists at $0.027 per million input tokens and $0.20 per million output tokens; Mercury 2 lists at $0.25 and $0.75.
Is Llama 3.2 1B or Mercury 2 better for coding?
Mercury 2 scores higher on coding benchmarks: 33.5 versus 21.1 in the Noometry coding category.
Which has the bigger context window?
Mercury 2 does, with 128K tokens against 60K.
How many benchmarks do Llama 3.2 1B and Mercury 2 share?
11 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Mercury 2 has 17.