Model comparison
Laguna M.1 vs Qwen3 8B
Qwen3 8B is the stronger model overall, scoring 33.7 to 32.5 on the Noometry Index.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in math, where Qwen3 8B leads 34.9 to 21.1.
- Laguna M.1 accepts more context: 262K tokens versus 131K.
Side by side
| Laguna M.1 | Qwen3 8B | |
|---|---|---|
| Provider | Poolside | Alibaba (Qwen) |
| Noometry Index | 32.5 | 33.7 |
| Released | 2026-04-28 | 2025-04 |
| Weights | Open | Open |
| Context window | 262K | 131K |
| Max output | 33K | 8K |
| Input $ / M tokens | — | $0.18 |
| Output $ / M tokens | — | $0.70 |
| Results tracked | 3 | 11 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Laguna M.1 leads
Laguna M.1: 36.6 (#204), Qwen3 8B: 34.0 (#248)
| Benchmark | Laguna M.1 | Qwen3 8B |
|---|---|---|
| LMArena WebDev | 1349 | — |
| SciCode | — | 22.6% |
Agentic & Tool Use Not comparable
Laguna M.1: —, Qwen3 8B: 30.2 (#78)
| Benchmark | Laguna M.1 | Qwen3 8B |
|---|---|---|
| Berkeley Function Calling Leaderboard | — | 42.6% |
Reasoning Laguna M.1 leads
Laguna M.1: 23.1 (#184), Qwen3 8B: 16.6 (#303)
| Benchmark | Laguna M.1 | Qwen3 8B |
|---|---|---|
| CritPt | — | 0% |
| Chess Puzzles | — | 5% |
| DTBench | — | 59.7% |
| LMCA | — | 8.8% |
| Surface Evolver Bench | 15.6% | — |
| Epoch Capabilities Index | — | 136.17 |
Math Qwen3 8B leads
Laguna M.1: 21.1 (#283), Qwen3 8B: 34.9 (#191)
| Benchmark | Laguna M.1 | Qwen3 8B |
|---|---|---|
| OTIS Mock AIME 2024-2025 | — | 56.1% |
| ProofBench | 0% | — |
Knowledge Not comparable
Laguna M.1: —, Qwen3 8B: 36.1 (#173)
| Benchmark | Laguna M.1 | Qwen3 8B |
|---|---|---|
| GPQA Diamond | — | 56.8% |
| Vectara Hallucination Rate | — | 4.8% |
Long Context Not comparable
Laguna M.1: —, Qwen3 8B: 37.9 (#210)
| Benchmark | Laguna M.1 | Qwen3 8B |
|---|---|---|
| Fiction.LiveBench | — | 62.1% |
Frequently asked questions
Is Laguna M.1 better than Qwen3 8B?
Qwen3 8B is the stronger model overall, scoring 33.7 to 32.5 on the Noometry Index.
Is Laguna M.1 or Qwen3 8B better for coding?
Laguna M.1 scores higher on coding benchmarks: 36.6 versus 34.0 in the Noometry coding category.
Which has the bigger context window?
Laguna M.1 does, with 262K tokens against 131K.
How many benchmarks do Laguna M.1 and Qwen3 8B share?
0 benchmarks have published results for both models. Laguna M.1 has 3 scored results on Noometry and Qwen3 8B has 11.