Model comparison

Laguna M.1 vs Llama 3.2 1B

Laguna M.1 is the stronger model overall, scoring 32.5 to 20.1 on the Noometry Index.

Last verified . 0 shared benchmarks.

Laguna M.1 Poolside

32.5

Rank #256 Reported

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Summary

  • The widest gap is in coding, where Laguna M.1 leads 36.6 to 21.1.
  • Laguna M.1 accepts more context: 262K tokens versus 60K.

Side by side

Laguna M.1 and Llama 3.2 1B specifications
Laguna M.1Llama 3.2 1B
ProviderPoolsideMeta
Noometry Index32.520.1
Released2026-04-282024-09-24
WeightsOpenOpen
Context window262K60K
Max output33K54K
Input $ / M tokens—$0.027
Output $ / M tokens—$0.20
Results tracked322

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Laguna M.1 leads

Laguna M.1: 36.6 (#204), Llama 3.2 1B: 21.1 (#338)

Coding benchmarks
BenchmarkLaguna M.1Llama 3.2 1B
LMArena WebDev1349—
BigCodeBench Instruct—8.2%
LMArena Coding—1070
BigCodeBench Complete—11.3%

Agentic & Tool Use Not comparable

Laguna M.1: —, Llama 3.2 1B: 14.6 (#150)

Agentic & Tool Use benchmarks
BenchmarkLaguna M.1Llama 3.2 1B
Berkeley Function Calling Leaderboard—10.8%
BALROG—6.6%

Reasoning Laguna M.1 leads

Laguna M.1: 23.1 (#184), Llama 3.2 1B: 16.2 (#308)

Reasoning benchmarks
BenchmarkLaguna M.1Llama 3.2 1B
Chess Puzzles—0%
LMArena Hard Prompts—1044
Surface Evolver Bench15.6%—
Epoch Capabilities Index—101.99

Math Laguna M.1 leads

Laguna M.1: 21.1 (#283), Llama 3.2 1B: 10.4 (#313)

Math benchmarks
BenchmarkLaguna M.1Llama 3.2 1B
OTIS Mock AIME 2024-2025—0.6%
ProofBench0%—
LMArena Math—1086

Knowledge Not comparable

Laguna M.1: —, Llama 3.2 1B: 7.2 (#312)

Knowledge benchmarks
BenchmarkLaguna M.1Llama 3.2 1B
GPQA Diamond—23.9%
LMArena Expert—1007

Multilingual Not comparable

Laguna M.1: —, Llama 3.2 1B: 23.8 (#292)

Multilingual benchmarks
BenchmarkLaguna M.1Llama 3.2 1B
LMArena Non-English—973
LMArena Chinese—959
LMArena German—1014
LMArena Russian—941

Instruction Following Not comparable

Laguna M.1: —, Llama 3.2 1B: 52.4 (#290)

Instruction Following benchmarks
BenchmarkLaguna M.1Llama 3.2 1B
LMArena Instruction Following—1031

Long Context Not comparable

Laguna M.1: —, Llama 3.2 1B: 31.9 (#274)

Long Context benchmarks
BenchmarkLaguna M.1Llama 3.2 1B
LMArena Longer Query—1050

Writing & Preference Not comparable

Laguna M.1: —, Llama 3.2 1B: 21.3 (#310)

Writing & Preference benchmarks
BenchmarkLaguna M.1Llama 3.2 1B
LMArena Text—1055
LMArena Creative Writing—1033
EQ-Bench Creative Writing—200
LMArena Multi-Turn—1030

Frequently asked questions

Is Laguna M.1 better than Llama 3.2 1B?

Laguna M.1 is the stronger model overall, scoring 32.5 to 20.1 on the Noometry Index.

Is Laguna M.1 or Llama 3.2 1B better for coding?

Laguna M.1 scores higher on coding benchmarks: 36.6 versus 21.1 in the Noometry coding category.

Which has the bigger context window?

Laguna M.1 does, with 262K tokens against 60K.

How many benchmarks do Laguna M.1 and Llama 3.2 1B share?

0 benchmarks have published results for both models. Laguna M.1 has 3 scored results on Noometry and Llama 3.2 1B has 22.

Related comparisons

Go deeper