Model comparison

Codestral vs Llama2 70b Steerlm Chat

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 30.6 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • Llama2 70b Steerlm Chat has downloadable open weights; the other is API-only.

Side by side

Codestral and Llama2 70b Steerlm Chat specifications
CodestralLlama2 70b Steerlm Chat
ProviderMistral AINVIDIA
Noometry Index30.631.8
Released2024-05-29—
WeightsProprietaryOpen
Context window256K—
Max output8K—
Input $ / M tokens$0.30—
Output $ / M tokens$0.90—
Results tracked79

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama2 70b Steerlm Chat leads

Codestral: 27.3 (#321), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkCodestralLlama2 70b Steerlm Chat
Aider Polyglot11.1%—
BigCodeBench Instruct41.8%—
LMArena Coding—1025
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Reasoning Too close to call

Codestral: 19.8 (#251), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkCodestralLlama2 70b Steerlm Chat
Kagi LLM Benchmark32.5%—
LMArena Hard Prompts—1047

Math Not comparable

Codestral: —, Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkCodestralLlama2 70b Steerlm Chat
LMArena Math—1072

Multilingual Not comparable

Codestral: —, Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkCodestralLlama2 70b Steerlm Chat
LMArena Non-English—1063

Instruction Following Not comparable

Codestral: —, Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkCodestralLlama2 70b Steerlm Chat
LMArena Instruction Following—1060

Long Context Not comparable

Codestral: —, Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkCodestralLlama2 70b Steerlm Chat
LMArena Longer Query—998

Writing & Preference Not comparable

Codestral: —, Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkCodestralLlama2 70b Steerlm Chat
LMArena Text—1098
LMArena Creative Writing—1091
LMArena Multi-Turn—1058

Frequently asked questions

Is Codestral better than Llama2 70b Steerlm Chat?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 30.6 on the Noometry Index.

Is Codestral or Llama2 70b Steerlm Chat better for coding?

Llama2 70b Steerlm Chat scores higher on coding benchmarks: 29.9 versus 27.3 in the Noometry coding category.

How many benchmarks do Codestral and Llama2 70b Steerlm Chat share?

0 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper