Model comparison

Laguna M.1 vs Mistral Large

Laguna M.1 and Mistral Large score almost the same on the Noometry Index (32.5 vs 31.9), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Laguna M.1 Poolside

32.5

Rank #256 Reported

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • The widest gap is in reasoning, where Laguna M.1 leads 23.1 to 15.8.
  • Laguna M.1 accepts more context: 262K tokens versus 131K.

Side by side

Laguna M.1 and Mistral Large specifications
Laguna M.1Mistral Large
ProviderPoolsideMistral AI
Noometry Index32.531.9
Released2026-04-282024-02-26
WeightsOpenOpen
Context window262K131K
Max output33K16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked351

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Laguna M.1 leads

Laguna M.1: 36.6 (#204), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkLaguna M.1Mistral Large
LMArena WebDev1349—
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
LMArena Coding—1277
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Laguna M.1: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkLaguna M.1Mistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Laguna M.1 leads

Laguna M.1: 23.1 (#184), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkLaguna M.1Mistral Large
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
LMArena Hard Prompts—1257
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Surface Evolver Bench15.6%—
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Laguna M.1 leads

Laguna M.1: 21.1 (#283), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkLaguna M.1Mistral Large
OTIS Mock AIME 2024-2025—8.5%
ProofBench0%—
Omni-MATH—28.1%
LiveBench Math—42.5%
LMArena Math—1262
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Not comparable

Laguna M.1: —, Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkLaguna M.1Mistral Large
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
LMArena Expert—1232
MMLU—80%

Multilingual Not comparable

Laguna M.1: —, Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkLaguna M.1Mistral Large
LMArena Non-English—1237
LMArena Chinese—1240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Not comparable

Laguna M.1: —, Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkLaguna M.1Mistral Large
LiveBench Instruction Following—67.9%
IFEval—87.7%
LMArena Instruction Following—1249

Long Context Not comparable

Laguna M.1: —, Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkLaguna M.1Mistral Large
LMArena Longer Query—1261

Writing & Preference Not comparable

Laguna M.1: —, Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkLaguna M.1Mistral Large
LMArena Text—1266
LMArena Creative Writing—1243
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LMArena Multi-Turn—1260
LiveBench Language—39.4%

Frequently asked questions

Is Laguna M.1 better than Mistral Large?

Laguna M.1 and Mistral Large score almost the same on the Noometry Index (32.5 vs 31.9), so choose on price, context window or the category you care about most.

Is Laguna M.1 or Mistral Large better for coding?

Laguna M.1 scores higher on coding benchmarks: 36.6 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Laguna M.1 does, with 262K tokens against 131K.

How many benchmarks do Laguna M.1 and Mistral Large share?

0 benchmarks have published results for both models. Laguna M.1 has 3 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper