Model comparison

Codestral vs Llama 3.1 Nemotron Ultra 253b v1

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 30.6 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Summary

  • The widest gap is in coding, where Llama 3.1 Nemotron Ultra 253b v1 leads 38.4 to 27.3.
  • Llama 3.1 Nemotron Ultra 253b v1 has downloadable open weights; the other is API-only.

Side by side

Codestral and Llama 3.1 Nemotron Ultra 253b v1 specifications
CodestralLlama 3.1 Nemotron Ultra 253b v1
ProviderMistral AINVIDIA
Noometry Index30.636.7
Released2024-05-29—
WeightsProprietaryOpen
Context window256K—
Max output8K—
Input $ / M tokens$0.30—
Output $ / M tokens$0.90—
Results tracked711

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron Ultra 253b v1 leads

Codestral: 27.3 (#321), Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177)

Coding benchmarks
BenchmarkCodestralLlama 3.1 Nemotron Ultra 253b v1
Aider Polyglot11.1%—
BigCodeBench Instruct41.8%—
LMArena Coding—1312
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Agentic & Tool Use Not comparable

Codestral: —, Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149)

Agentic & Tool Use benchmarks
BenchmarkCodestralLlama 3.1 Nemotron Ultra 253b v1
Berkeley Function Calling Leaderboard—10%

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

Codestral: 19.8 (#251), Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134)

Reasoning benchmarks
BenchmarkCodestralLlama 3.1 Nemotron Ultra 253b v1
Kagi LLM Benchmark32.5%—
LMArena Hard Prompts—1316

Math Not comparable

Codestral: —, Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152)

Math benchmarks
BenchmarkCodestralLlama 3.1 Nemotron Ultra 253b v1
LMArena Math—1360

Multilingual Not comparable

Codestral: —, Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187)

Multilingual benchmarks
BenchmarkCodestralLlama 3.1 Nemotron Ultra 253b v1
LMArena Non-English—1282
LMArena Russian—1284

Instruction Following Not comparable

Codestral: —, Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178)

Instruction Following benchmarks
BenchmarkCodestralLlama 3.1 Nemotron Ultra 253b v1
LMArena Instruction Following—1308

Long Context Not comparable

Codestral: —, Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177)

Long Context benchmarks
BenchmarkCodestralLlama 3.1 Nemotron Ultra 253b v1
LMArena Longer Query—1299

Writing & Preference Not comparable

Codestral: —, Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175)

Writing & Preference benchmarks
BenchmarkCodestralLlama 3.1 Nemotron Ultra 253b v1
LMArena Text—1320
LMArena Creative Writing—1314
LMArena Multi-Turn—1317

Frequently asked questions

Is Codestral better than Llama 3.1 Nemotron Ultra 253b v1?

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 30.6 on the Noometry Index.

Is Codestral or Llama 3.1 Nemotron Ultra 253b v1 better for coding?

Llama 3.1 Nemotron Ultra 253b v1 scores higher on coding benchmarks: 38.4 versus 27.3 in the Noometry coding category.

How many benchmarks do Codestral and Llama 3.1 Nemotron Ultra 253b v1 share?

0 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Llama 3.1 Nemotron Ultra 253b v1 has 11.

Related comparisons

Go deeper