Model comparison

Llama-3.3-70B-Instruct vs Mistral Large

Mistral Large is the stronger model overall, scoring 31.9 to 30.6 on the Noometry Index. Llama-3.3-70B-Instruct costs 19× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Last verified . 40 shared benchmarks.

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 40 benchmarks with published results for both. Llama-3.3-70B-Instruct scores higher in 3 categories and Mistral Large in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Mistral Large leads 38.3 to 26.4.
  • The biggest single-benchmark swing is BigCodeBench Complete: 57.5% for Llama-3.3-70B-Instruct and 38.3% for Mistral Large.
  • Llama-3.3-70B-Instruct is cheaper at $0.10 / $0.32 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Mistral Large accepts more context: 131K tokens versus 128K.

Side by side

Llama-3.3-70B-Instruct and Mistral Large specifications
Llama-3.3-70B-InstructMistral Large
ProviderMetaMistral AI
Noometry Index30.631.9
Released2024-12-062024-02-26
WeightsOpenOpen
Context window128K131K
Max output4K16K
Input $ / M tokens$0.10$2
Output $ / M tokens$0.32$6
Results tracked4351

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Llama-3.3-70B-Instruct: 31.0 (#290), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkLlama-3.3-70B-InstructMistral Large
SciCode26%36.2%
BigCodeBench Instruct46.9%30%
LiveBench Coding36.6%47.1%
LMArena Coding12681277
BigCodeBench Complete57.5%38.3%
WeirdML14.4%—
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Mistral Large leads

Llama-3.3-70B-Instruct: 25.8 (#105), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkLlama-3.3-70B-InstructMistral Large
Berkeley Function Calling Leaderboard31.9%38.4%
BALROG23%—

Reasoning Mistral Large leads

Llama-3.3-70B-Instruct: 14.1 (#327), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkLlama-3.3-70B-InstructMistral Large
SimpleBench19.9%22.5%
CritPt0%0%
LiveBench Reasoning50.8%43.5%
LMArena Hard Prompts12571257
DTBench59.5%65.1%
LiveBench Data Analysis49.5%50.1%
LMCA17.5%16.7%
Epoch Capabilities Index127.33128.52
ForecastBench58.657.1
LiveBench50.2%48.4%

Math Mistral Large leads

Llama-3.3-70B-Instruct: 15.3 (#298), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkLlama-3.3-70B-InstructMistral Large
OTIS Mock AIME 2024-20255.1%8.5%
LiveBench Math42.2%42.5%
LMArena Math12671262
MATH Level 541.6%50.3%
Omni-MATH—28.1%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Too close to call

Llama-3.3-70B-Instruct: 30.6 (#226), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkLlama-3.3-70B-InstructMistral Large
GPQA Diamond47.4%51.3%
Confabulations22.8%21.4%
Vectara Hallucination Rate4.1%4.5%
LMArena Expert12251232
MMLU86.3%80%
MMLU-Pro—59.9%
GPQA (HELM)—43.5%

Multilingual Too close to call

Llama-3.3-70B-Instruct: 39.9 (#220), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkLlama-3.3-70B-InstructMistral Large
LMArena Non-English12361237
LMArena Chinese12171240
LMArena French12811325
LMArena German12511254
LMArena Japanese11501188
LMArena Korean11431202
LMArena Russian12521257
LMArena Spanish12701268

Instruction Following Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 71.1 (#157), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkLlama-3.3-70B-InstructMistral Large
LiveBench Instruction Following82.7%67.9%
LMArena Instruction Following12421249
IFEval—87.7%

Long Context Mistral Large leads

Llama-3.3-70B-Instruct: 26.4 (#295), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkLlama-3.3-70B-InstructMistral Large
LMArena Longer Query12561261
Fiction.LiveBench33.3%—

Writing & Preference Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 47.6 (#207), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkLlama-3.3-70B-InstructMistral Large
LMArena Text12741266
LMArena Creative Writing12501243
LMArena Multi-Turn12801260
LiveBench Language39.2%39.4%
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%

Frequently asked questions

Is Llama-3.3-70B-Instruct better than Mistral Large?

Mistral Large is the stronger model overall, scoring 31.9 to 30.6 on the Noometry Index. Llama-3.3-70B-Instruct costs 19× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Which is cheaper, Llama-3.3-70B-Instruct or Mistral Large?

Llama-3.3-70B-Instruct is cheaper. It lists at $0.10 per million input tokens and $0.32 per million output tokens; Mistral Large lists at $2 and $6.

Is Llama-3.3-70B-Instruct or Mistral Large better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 31.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Large does, with 131K tokens against 128K.

How many benchmarks do Llama-3.3-70B-Instruct and Mistral Large share?

40 benchmarks have published results for both models. Llama-3.3-70B-Instruct has 43 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper