Model comparison

Mistral Large vs Nemotron 3.5 Lightning

Nemotron 3.5 Lightning is the stronger model overall, scoring 40.0 to 31.9 on the Noometry Index.

Last verified . 18 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Nemotron 3.5 Lightning NVIDIA

40.0

Rank #155 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Mistral Large scores higher in 0 categories and Nemotron 3.5 Lightning in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Nemotron 3.5 Lightning leads 37.5 to 18.2.
  • Nemotron 3.5 Lightning is cheaper at $0.05 / $0.20 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Nemotron 3.5 Lightning accepts more context: 262K tokens versus 131K.

Side by side

Mistral Large and Nemotron 3.5 Lightning specifications
Mistral LargeNemotron 3.5 Lightning
ProviderMistral AINVIDIA
Noometry Index31.940.0
Released2024-02-262026-08-11
WeightsOpenOpen
Context window131K262K
Max output16K262K
Input $ / M tokens$2$0.05
Output $ / M tokens$6$0.20
Results tracked5118

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nemotron 3.5 Lightning leads

Mistral Large: 34.3 (#240), Nemotron 3.5 Lightning: 40.4 (#141)

Coding benchmarks
BenchmarkMistral LargeNemotron 3.5 Lightning
LMArena Coding12771375
SciCode36.2%—
BigCodeBench Instruct30%—
LiveBench Coding47.1%—
BigCodeBench Complete38.3%—
ALE-Bench264.7—
HumanEval+62.2%—
MBPP+59.5%—

Agentic & Tool Use Not comparable

Mistral Large: 28.6 (#89), Nemotron 3.5 Lightning: —

Agentic & Tool Use benchmarks
BenchmarkMistral LargeNemotron 3.5 Lightning
Berkeley Function Calling Leaderboard38.4%—

Reasoning Nemotron 3.5 Lightning leads

Mistral Large: 15.8 (#310), Nemotron 3.5 Lightning: 26.8 (#127)

Reasoning benchmarks
BenchmarkMistral LargeNemotron 3.5 Lightning
LMArena Hard Prompts12571337
SimpleBench22.5%—
CritPt0%—
LiveBench Reasoning43.5%—
DTBench65.1%—
LiveBench Data Analysis50.1%—
LMCA16.7%—
Epoch Capabilities Index128.52—
ForecastBench57.1—
LiveBench48.4%—

Math Nemotron 3.5 Lightning leads

Mistral Large: 18.2 (#291), Nemotron 3.5 Lightning: 37.5 (#155)

Math benchmarks
BenchmarkMistral LargeNemotron 3.5 Lightning
LMArena Math12621359
OTIS Mock AIME 2024-20258.5%—
Omni-MATH28.1%—
LiveBench Math42.5%—
MATH Level 550.3%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Nemotron 3.5 Lightning leads

Mistral Large: 30.1 (#230), Nemotron 3.5 Lightning: 37.5 (#154)

Knowledge benchmarks
BenchmarkMistral LargeNemotron 3.5 Lightning
LMArena Expert12321356
GPQA Diamond51.3%—
MMLU-Pro59.9%—
Confabulations21.4%—
Vectara Hallucination Rate4.5%—
GPQA (HELM)43.5%—
MMLU80%—

Multilingual Nemotron 3.5 Lightning leads

Mistral Large: 40.0 (#219), Nemotron 3.5 Lightning: 44.0 (#180)

Multilingual benchmarks
BenchmarkMistral LargeNemotron 3.5 Lightning
LMArena Non-English12371295
LMArena Chinese12401359
LMArena French13251366
LMArena German12541282
LMArena Japanese11881206
LMArena Korean12021238
LMArena Russian12571253
LMArena Spanish12681345

Instruction Following Nemotron 3.5 Lightning leads

Mistral Large: 67.9 (#191), Nemotron 3.5 Lightning: 69.6 (#170)

Instruction Following benchmarks
BenchmarkMistral LargeNemotron 3.5 Lightning
LMArena Instruction Following12491318
LiveBench Instruction Following67.9%—
IFEval87.7%—

Long Context Nemotron 3.5 Lightning leads

Mistral Large: 38.3 (#199), Nemotron 3.5 Lightning: 39.9 (#165)

Long Context benchmarks
BenchmarkMistral LargeNemotron 3.5 Lightning
LMArena Longer Query12611314

Writing & Preference Nemotron 3.5 Lightning leads

Mistral Large: 40.7 (#242), Nemotron 3.5 Lightning: 48.5 (#201)

Writing & Preference benchmarks
BenchmarkMistral LargeNemotron 3.5 Lightning
LMArena Text12661327
LMArena Creative Writing12431254
EQ-Bench Creative Writing9851280
LMArena Multi-Turn12601328
Short-Story Creative Writing69%—
WildBench80.1%—
LiveBench Language39.4%—

Frequently asked questions

Is Mistral Large better than Nemotron 3.5 Lightning?

Nemotron 3.5 Lightning is the stronger model overall, scoring 40.0 to 31.9 on the Noometry Index.

Which is cheaper, Mistral Large or Nemotron 3.5 Lightning?

Nemotron 3.5 Lightning is cheaper. It lists at $0.05 per million input tokens and $0.20 per million output tokens; Mistral Large lists at $2 and $6.

Is Mistral Large or Nemotron 3.5 Lightning better for coding?

Nemotron 3.5 Lightning scores higher on coding benchmarks: 40.4 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Nemotron 3.5 Lightning does, with 262K tokens against 131K.

How many benchmarks do Mistral Large and Nemotron 3.5 Lightning share?

18 benchmarks have published results for both models. Mistral Large has 51 scored results on Noometry and Nemotron 3.5 Lightning has 18.

Related comparisons

Go deeper