Model comparison

Codestral vs Llama-3.3-70B-Instruct

Codestral and Llama-3.3-70B-Instruct score almost the same on the Noometry Index (30.6 vs 30.6), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Codestral scores higher in 1 category and Llama-3.3-70B-Instruct in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Codestral leads 19.8 to 14.1.
  • The biggest single-benchmark swing is BigCodeBench Instruct: 41.8% for Codestral and 46.9% for Llama-3.3-70B-Instruct.
  • Llama-3.3-70B-Instruct is cheaper at $0.10 / $0.32 per million input/output tokens, against $0.30 / $0.90 for Codestral.
  • Codestral accepts more context: 256K tokens versus 128K.
  • Llama-3.3-70B-Instruct has downloadable open weights; the other is API-only.

Side by side

Codestral and Llama-3.3-70B-Instruct specifications
CodestralLlama-3.3-70B-Instruct
ProviderMistral AIMeta
Noometry Index30.630.6
Released2024-05-292024-12-06
WeightsProprietaryOpen
Context window256K128K
Max output8K4K
Input $ / M tokens$0.30$0.10
Output $ / M tokens$0.90$0.32
Results tracked743

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama-3.3-70B-Instruct leads

Codestral: 27.3 (#321), Llama-3.3-70B-Instruct: 31.0 (#290)

Coding benchmarks
BenchmarkCodestralLlama-3.3-70B-Instruct
BigCodeBench Instruct41.8%46.9%
BigCodeBench Complete52.5%57.5%
Aider Polyglot11.1%—
SciCode—26%
WeirdML—14.4%
LiveBench Coding—36.6%
LMArena Coding—1268
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Agentic & Tool Use Not comparable

Codestral: —, Llama-3.3-70B-Instruct: 25.8 (#105)

Agentic & Tool Use benchmarks
BenchmarkCodestralLlama-3.3-70B-Instruct
Berkeley Function Calling Leaderboard—31.9%
BALROG—23%

Reasoning Codestral leads

Codestral: 19.8 (#251), Llama-3.3-70B-Instruct: 14.1 (#327)

Reasoning benchmarks
BenchmarkCodestralLlama-3.3-70B-Instruct
SimpleBench—19.9%
Kagi LLM Benchmark32.5%—
CritPt—0%
LiveBench Reasoning—50.8%
LMArena Hard Prompts—1257
DTBench—59.5%
LiveBench Data Analysis—49.5%
LMCA—17.5%
Epoch Capabilities Index—127.33
ForecastBench—58.6
LiveBench—50.2%

Math Not comparable

Codestral: —, Llama-3.3-70B-Instruct: 15.3 (#298)

Math benchmarks
BenchmarkCodestralLlama-3.3-70B-Instruct
OTIS Mock AIME 2024-2025—5.1%
LiveBench Math—42.2%
LMArena Math—1267
MATH Level 5—41.6%

Knowledge Not comparable

Codestral: —, Llama-3.3-70B-Instruct: 30.6 (#226)

Knowledge benchmarks
BenchmarkCodestralLlama-3.3-70B-Instruct
GPQA Diamond—47.4%
Confabulations—22.8%
Vectara Hallucination Rate—4.1%
LMArena Expert—1225
MMLU—86.3%

Multilingual Not comparable

Codestral: —, Llama-3.3-70B-Instruct: 39.9 (#220)

Multilingual benchmarks
BenchmarkCodestralLlama-3.3-70B-Instruct
LMArena Non-English—1236
LMArena Chinese—1217
LMArena French—1281
LMArena German—1251
LMArena Japanese—1150
LMArena Korean—1143
LMArena Russian—1252
LMArena Spanish—1270

Instruction Following Not comparable

Codestral: —, Llama-3.3-70B-Instruct: 71.1 (#157)

Instruction Following benchmarks
BenchmarkCodestralLlama-3.3-70B-Instruct
LiveBench Instruction Following—82.7%
LMArena Instruction Following—1242

Long Context Not comparable

Codestral: —, Llama-3.3-70B-Instruct: 26.4 (#295)

Long Context benchmarks
BenchmarkCodestralLlama-3.3-70B-Instruct
Fiction.LiveBench—33.3%
LMArena Longer Query—1256

Writing & Preference Not comparable

Codestral: —, Llama-3.3-70B-Instruct: 47.6 (#207)

Writing & Preference benchmarks
BenchmarkCodestralLlama-3.3-70B-Instruct
LMArena Text—1274
LMArena Creative Writing—1250
LMArena Multi-Turn—1280
LiveBench Language—39.2%

Frequently asked questions

Is Codestral better than Llama-3.3-70B-Instruct?

Codestral and Llama-3.3-70B-Instruct score almost the same on the Noometry Index (30.6 vs 30.6), so choose on price, context window or the category you care about most.

Which is cheaper, Codestral or Llama-3.3-70B-Instruct?

Llama-3.3-70B-Instruct is cheaper. It lists at $0.10 per million input tokens and $0.32 per million output tokens; Codestral lists at $0.30 and $0.90.

Is Codestral or Llama-3.3-70B-Instruct better for coding?

Llama-3.3-70B-Instruct scores higher on coding benchmarks: 31.0 versus 27.3 in the Noometry coding category.

Which has the bigger context window?

Codestral does, with 256K tokens against 128K.

How many benchmarks do Codestral and Llama-3.3-70B-Instruct share?

2 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Llama-3.3-70B-Instruct has 43.

Related comparisons

Go deeper