Model comparison

Laguna M.1 vs o3

o3 is the stronger model overall, scoring 47.5 to 32.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Laguna M.1 Poolside

32.5

Rank #256 Reported

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • The widest gap is in math, where o3 leads 50.2 to 21.1.
  • Laguna M.1 accepts more context: 262K tokens versus 200K.
  • Laguna M.1 has downloadable open weights; the other is API-only.

Side by side

Laguna M.1 and o3 specifications
Laguna M.1o3
ProviderPoolsideOpenAI
Noometry Index32.547.5
Released2026-04-282025-04-16
WeightsOpenProprietary
Context window262K200K
Max output33K100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked363

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Laguna M.1: 36.6 (#204), o3: 46.8 (#64)

Coding benchmarks
BenchmarkLaguna M.1o3
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
LMArena WebDev1349—
GSO—8.8%
WeirdML—52.4%
LMArena Coding—1408
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Laguna M.1: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkLaguna M.1o3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Laguna M.1: 23.1 (#184), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkLaguna M.1o3
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
LMArena Hard Prompts—1402
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
Surface Evolver Bench15.6%—
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math o3 leads

Laguna M.1: 21.1 (#283), o3: 50.2 (#58)

Math benchmarks
BenchmarkLaguna M.1o3
FrontierMath (Tiers 1-3)—33.3%
OTIS Mock AIME 2024-2025—84.4%
ProofBench0%—
Omni-MATH—71.4%
LMArena Math—1426
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Not comparable

Laguna M.1: —, o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkLaguna M.1o3
GPQA Diamond—81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%
LMArena Expert—1402

Multimodal Not comparable

Laguna M.1: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkLaguna M.1o3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual Not comparable

Laguna M.1: —, o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkLaguna M.1o3
LMArena Non-English—1401
LMArena Chinese—1437
LMArena French—1430
LMArena German—1420
LMArena Japanese—1403
LMArena Korean—1370
LMArena Russian—1406
LMArena Spanish—1395

Instruction Following Not comparable

Laguna M.1: —, o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkLaguna M.1o3
IFEval—86.9%
LMArena Instruction Following—1368

Long Context Not comparable

Laguna M.1: —, o3: 53.3 (#6)

Long Context benchmarks
BenchmarkLaguna M.1o3
Fiction.LiveBench—88.9%
CL-bench—17.8%
LMArena Longer Query—1372

Writing & Preference Not comparable

Laguna M.1: —, o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkLaguna M.1o3
LMArena Text—1410
LMArena Creative Writing—1359
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%
LMArena Multi-Turn—1405

Frequently asked questions

Is Laguna M.1 better than o3?

o3 is the stronger model overall, scoring 47.5 to 32.5 on the Noometry Index.

Is Laguna M.1 or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 36.6 in the Noometry coding category.

Which has the bigger context window?

Laguna M.1 does, with 262K tokens against 200K.

How many benchmarks do Laguna M.1 and o3 share?

0 benchmarks have published results for both models. Laguna M.1 has 3 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper