Model comparison

Llama 4 Maverick vs Mistral Small 3.1

Llama 4 Maverick and Mistral Small 3.1 score almost the same on the Noometry Index (30.9 vs 31.7), so choose on price, context window or the category you care about most.

Last verified . 27 shared benchmarks.

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Llama 4 Maverick scores higher in 5 categories and Mistral Small 3.1 in 4 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Mistral Small 3.1 leads 38.3 to 26.6.
  • The biggest single-benchmark swing is GPQA (HELM): 65% for Llama 4 Maverick and 39.2% for Mistral Small 3.1.
  • Llama 4 Maverick is cheaper at $0.19 / $0.65 per million input/output tokens, against $0.35 / $0.56 for Mistral Small 3.1.

Side by side

Llama 4 Maverick and Mistral Small 3.1 specifications
Llama 4 MaverickMistral Small 3.1
ProviderMetaMistral AI
Noometry Index30.931.7
Released2025-04-052025-03-17
WeightsOpenOpen
Context window128K128K
Max output4K102K
Input $ / M tokens$0.19$0.35
Output $ / M tokens$0.65$0.56
Results tracked5428

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3.1 leads

Llama 4 Maverick: 26.6 (#324), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.1
LMArena Coding13021309
SWE-bench Verified (bash only)21%—
Aider Polyglot15.6%—
SciCode33.1%—
WeirdML24.5%—
BigCodeBench Instruct49.7%—
BigCodeBench Complete61.4%—
ALE-Bench172.97—

Agentic & Tool Use Not comparable

Llama 4 Maverick: 28.2 (#91), Mistral Small 3.1: —

Agentic & Tool Use benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.1
Berkeley Function Calling Leaderboard37.3%—

Reasoning Mistral Small 3.1 leads

Llama 4 Maverick: 10.1 (#342), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.1
LMArena Hard Prompts12811278
Epoch Capabilities Index132.2127.48
ARC-AGI-20%—
SimpleBench27.7%—
Kagi LLM Benchmark55.9%—
NYT Connections (extended)8%—
ARC-AGI-14.4%—
CritPt0%—
Chess Puzzles—1%
EnigmaEval0.6%—
DTBench61.9%—
LMCA15.9%—
ForecastBench57.5—

Math Llama 4 Maverick leads

Llama 4 Maverick: 26.0 (#262), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.1
OTIS Mock AIME 2024-202520.6%3.9%
Omni-MATH42.2%24.8%
LMArena Math12991262
MATH Level 573%—
FrontierMath (Feb 2025 set)0.7%—

Knowledge Llama 4 Maverick leads

Llama 4 Maverick: 33.4 (#204), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.1
GPQA Diamond67%41.9%
MMLU-Pro81%61%
GPQA (HELM)65%39.2%
LMArena Expert12591257
Humanity's Last Exam5.7%—
Confabulations22.6%—
Vectara Hallucination Rate8.2%—

Multimodal Mistral Small 3.1 leads

Llama 4 Maverick: 31.6 (#105), Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.1
LMArena Vision11421136
GeoBench52%—
SpatialViz-Bench31.8%—

Multilingual Llama 4 Maverick leads

Llama 4 Maverick: 42.2 (#195), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.1
LMArena Non-English12691255
LMArena Chinese12771253
LMArena French12591273
LMArena German12911266
LMArena Japanese12071208
LMArena Korean12031206
LMArena Russian12861263
LMArena Spanish12931283

Instruction Following Llama 4 Maverick leads

Llama 4 Maverick: 71.7 (#146), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.1
IFEval90.8%75%
LMArena Instruction Following12671264

Long Context Mistral Small 3.1 leads

Llama 4 Maverick: 31.4 (#279), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.1
LMArena Longer Query12801299
Fiction.LiveBench46.2%—

Writing & Preference Llama 4 Maverick leads

Llama 4 Maverick: 38.8 (#252), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.1
LMArena Text12871277
LMArena Creative Writing12671253
EQ-Bench Creative Writing860761
WildBench80%78.8%
LMArena Multi-Turn12891270
Short-Story Creative Writing62%—

Frequently asked questions

Is Llama 4 Maverick better than Mistral Small 3.1?

Llama 4 Maverick and Mistral Small 3.1 score almost the same on the Noometry Index (30.9 vs 31.7), so choose on price, context window or the category you care about most.

Which is cheaper, Llama 4 Maverick or Mistral Small 3.1?

Llama 4 Maverick is cheaper. It lists at $0.19 per million input tokens and $0.65 per million output tokens; Mistral Small 3.1 lists at $0.35 and $0.56.

Is Llama 4 Maverick or Mistral Small 3.1 better for coding?

Mistral Small 3.1 scores higher on coding benchmarks: 38.3 versus 26.6 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Llama 4 Maverick and Mistral Small 3.1 share?

27 benchmarks have published results for both models. Llama 4 Maverick has 54 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper