Model comparison

Llama 4 Maverick vs Mistral Nemo

Llama 4 Maverick is the stronger model overall, scoring 30.9 to 26.4 on the Noometry Index. Mistral Nemo costs 2.0× less per token, which makes it the better buy when Llama 4 Maverick's lead doesn't matter for your workload.

Last verified . 6 shared benchmarks.

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Llama 4 Maverick scores higher in 4 categories and Mistral Nemo in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama 4 Maverick leads 33.4 to 12.3.
  • The biggest single-benchmark swing is MATH Level 5: 73% for Llama 4 Maverick and 10.8% for Mistral Nemo.
  • Mistral Nemo is cheaper at $0.15 / $0.15 per million input/output tokens, against $0.19 / $0.65 for Llama 4 Maverick.

Side by side

Llama 4 Maverick and Mistral Nemo specifications
Llama 4 MaverickMistral Nemo
ProviderMetaMistral AI
Noometry Index30.926.4
Released2025-04-052024-07-01
WeightsOpenOpen
Context window128K128K
Max output4K128K
Input $ / M tokens$0.19$0.15
Output $ / M tokens$0.65$0.15
Results tracked5410

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 4 Maverick: 26.6 (#324), Mistral Nemo: —

Coding benchmarks
BenchmarkLlama 4 MaverickMistral Nemo
SWE-bench Verified (bash only)21%—
Aider Polyglot15.6%—
SciCode33.1%—
WeirdML24.5%—
BigCodeBench Instruct49.7%—
LMArena Coding1302—
BigCodeBench Complete61.4%—
ALE-Bench172.97—

Agentic & Tool Use Llama 4 Maverick leads

Llama 4 Maverick: 28.2 (#91), Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkLlama 4 MaverickMistral Nemo
Berkeley Function Calling Leaderboard37.3%27.6%
BALROG—17.6%

Reasoning Mistral Nemo leads

Llama 4 Maverick: 10.1 (#342), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkLlama 4 MaverickMistral Nemo
DTBench61.9%48.6%
Epoch Capabilities Index132.2118.68
ARC-AGI-20%—
SimpleBench27.7%—
Kagi LLM Benchmark55.9%—
NYT Connections (extended)8%—
ARC-AGI-14.4%—
CritPt0%—
EnigmaEval0.6%—
LMArena Hard Prompts1281—
LMCA15.9%—
ForecastBench57.5—
PIQA—83.5%

Math Too close to call

Llama 4 Maverick: 26.0 (#262), Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkLlama 4 MaverickMistral Nemo
MATH Level 573%10.8%
OTIS Mock AIME 2024-202520.6%—
Omni-MATH42.2%—
LMArena Math1299—
FrontierMath (Feb 2025 set)0.7%—
GSM8K—84.2%

Knowledge Llama 4 Maverick leads

Llama 4 Maverick: 33.4 (#204), Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkLlama 4 MaverickMistral Nemo
GPQA Diamond67%29.9%
Humanity's Last Exam5.7%—
MMLU-Pro81%—
Confabulations22.6%—
Vectara Hallucination Rate8.2%—
GPQA (HELM)65%—
LMArena Expert1259—
BoolQ—82.5%

Multimodal Not comparable

Llama 4 Maverick: 31.6 (#105), Mistral Nemo: —

Multimodal benchmarks
BenchmarkLlama 4 MaverickMistral Nemo
LMArena Vision1142—
GeoBench52%—
SpatialViz-Bench31.8%—

Multilingual Not comparable

Llama 4 Maverick: 42.2 (#195), Mistral Nemo: —

Multilingual benchmarks
BenchmarkLlama 4 MaverickMistral Nemo
LMArena Non-English1269—
LMArena Chinese1277—
LMArena French1259—
LMArena German1291—
LMArena Japanese1207—
LMArena Korean1203—
LMArena Russian1286—
LMArena Spanish1293—

Instruction Following Not comparable

Llama 4 Maverick: 71.7 (#146), Mistral Nemo: —

Instruction Following benchmarks
BenchmarkLlama 4 MaverickMistral Nemo
IFEval90.8%—
LMArena Instruction Following1267—

Long Context Not comparable

Llama 4 Maverick: 31.4 (#279), Mistral Nemo: —

Long Context benchmarks
BenchmarkLlama 4 MaverickMistral Nemo
Fiction.LiveBench46.2%—
LMArena Longer Query1280—

Writing & Preference Llama 4 Maverick leads

Llama 4 Maverick: 38.8 (#252), Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkLlama 4 MaverickMistral Nemo
EQ-Bench Creative Writing860881
LMArena Text1287—
LMArena Creative Writing1267—
Short-Story Creative Writing62%—
WildBench80%—
LMArena Multi-Turn1289—

Frequently asked questions

Is Llama 4 Maverick better than Mistral Nemo?

Llama 4 Maverick is the stronger model overall, scoring 30.9 to 26.4 on the Noometry Index. Mistral Nemo costs 2.0× less per token, which makes it the better buy when Llama 4 Maverick's lead doesn't matter for your workload.

Which is cheaper, Llama 4 Maverick or Mistral Nemo?

Mistral Nemo is cheaper. It lists at $0.15 per million input tokens and $0.15 per million output tokens; Llama 4 Maverick lists at $0.19 and $0.65.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Llama 4 Maverick and Mistral Nemo share?

6 benchmarks have published results for both models. Llama 4 Maverick has 54 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper