Model comparison

Laguna M.1 vs Qwen3 8B

Qwen3 8B is the stronger model overall, scoring 33.7 to 32.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Laguna M.1 Poolside

32.5

Rank #256 Reported

Qwen3 8B Alibaba (Qwen)

33.7

Rank #238 Confirmed

Summary

  • The widest gap is in math, where Qwen3 8B leads 34.9 to 21.1.
  • Laguna M.1 accepts more context: 262K tokens versus 131K.

Side by side

Laguna M.1 and Qwen3 8B specifications
Laguna M.1Qwen3 8B
ProviderPoolsideAlibaba (Qwen)
Noometry Index32.533.7
Released2026-04-282025-04
WeightsOpenOpen
Context window262K131K
Max output33K8K
Input $ / M tokens—$0.18
Output $ / M tokens—$0.70
Results tracked311

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Laguna M.1 leads

Laguna M.1: 36.6 (#204), Qwen3 8B: 34.0 (#248)

Coding benchmarks
BenchmarkLaguna M.1Qwen3 8B
LMArena WebDev1349—
SciCode—22.6%

Agentic & Tool Use Not comparable

Laguna M.1: —, Qwen3 8B: 30.2 (#78)

Agentic & Tool Use benchmarks
BenchmarkLaguna M.1Qwen3 8B
Berkeley Function Calling Leaderboard—42.6%

Reasoning Laguna M.1 leads

Laguna M.1: 23.1 (#184), Qwen3 8B: 16.6 (#303)

Reasoning benchmarks
BenchmarkLaguna M.1Qwen3 8B
CritPt—0%
Chess Puzzles—5%
DTBench—59.7%
LMCA—8.8%
Surface Evolver Bench15.6%—
Epoch Capabilities Index—136.17

Math Qwen3 8B leads

Laguna M.1: 21.1 (#283), Qwen3 8B: 34.9 (#191)

Math benchmarks
BenchmarkLaguna M.1Qwen3 8B
OTIS Mock AIME 2024-2025—56.1%
ProofBench0%—

Knowledge Not comparable

Laguna M.1: —, Qwen3 8B: 36.1 (#173)

Knowledge benchmarks
BenchmarkLaguna M.1Qwen3 8B
GPQA Diamond—56.8%
Vectara Hallucination Rate—4.8%

Long Context Not comparable

Laguna M.1: —, Qwen3 8B: 37.9 (#210)

Long Context benchmarks
BenchmarkLaguna M.1Qwen3 8B
Fiction.LiveBench—62.1%

Frequently asked questions

Is Laguna M.1 better than Qwen3 8B?

Qwen3 8B is the stronger model overall, scoring 33.7 to 32.5 on the Noometry Index.

Is Laguna M.1 or Qwen3 8B better for coding?

Laguna M.1 scores higher on coding benchmarks: 36.6 versus 34.0 in the Noometry coding category.

Which has the bigger context window?

Laguna M.1 does, with 262K tokens against 131K.

How many benchmarks do Laguna M.1 and Qwen3 8B share?

0 benchmarks have published results for both models. Laguna M.1 has 3 scored results on Noometry and Qwen3 8B has 11.

Related comparisons

Go deeper