Model comparison

Codellama 70b Instruct vs Mixtral 8x7B

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 27.1 on the Noometry Index.

Last verified . 5 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Codellama 70b Instruct scores higher in 3 categories and Mixtral 8x7B in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Codellama 70b Instruct leads 37.6 to 32.8.

Side by side

Codellama 70b Instruct and Mixtral 8x7B specifications
Codellama 70b InstructMixtral 8x7B
ProviderMetaMistral AI
Noometry Index33.727.1
Released—2023-12-11
WeightsOpenOpen
Context window—32K
Max output—32K
Input $ / M tokens—$0.70
Output $ / M tokens—$0.70
Results tracked738

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 70b Instruct leads

Codellama 70b Instruct: 37.6 (#193), Mixtral 8x7B: 32.8 (#269)

Coding benchmarks
BenchmarkCodellama 70b InstructMixtral 8x7B
HumanEval+65.9%39.6%
BigCodeBench Instruct40.7%—
LMArena Coding—1126
BigCodeBench Complete49.6%—
MBPP+—49.7%

Reasoning Codellama 70b Instruct leads

Codellama 70b Instruct: 20.1 (#242), Mixtral 8x7B: 18.2 (#285)

Reasoning benchmarks
BenchmarkCodellama 70b InstructMixtral 8x7B
LMArena Hard Prompts10521115
DTBench—49.6%
Adversarial NLI—55.2%
Epoch Capabilities Index—118.47
ForecastBench—56.3
HellaSwag—86.7%
PIQA—83.6%
WinoGrande—77.2%

Math Not comparable

Codellama 70b Instruct: —, Mixtral 8x7B: 18.8 (#289)

Math benchmarks
BenchmarkCodellama 70b InstructMixtral 8x7B
Omni-MATH—10.5%
LMArena Math—1147
MATH Level 5—10%
GSM8K—74.4%

Knowledge Not comparable

Codellama 70b Instruct: —, Mixtral 8x7B: 11.0 (#301)

Knowledge benchmarks
BenchmarkCodellama 70b InstructMixtral 8x7B
GPQA Diamond—30.6%
MMLU-Pro—33.5%
GPQA (HELM)—29.6%
LMArena Expert—1088
ARC (AI2) Challenge—87.3%
MMLU—70.6%
OpenBookQA—85.8%
TriviaQA—82.2%

Multilingual Mixtral 8x7B leads

Codellama 70b Instruct: 24.8 (#288), Mixtral 8x7B: 29.6 (#266)

Multilingual benchmarks
BenchmarkCodellama 70b InstructMixtral 8x7B
LMArena Non-English9921077
LMArena Chinese—1055
LMArena French—1166
LMArena German—1114
LMArena Japanese—931
LMArena Korean—968
LMArena Russian—1090
LMArena Spanish—1111

Instruction Following Too close to call

Codellama 70b Instruct: 51.9 (#293), Mixtral 8x7B: 51.0 (#297)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructMixtral 8x7B
LMArena Instruction Following10241109
IFEval—57.5%

Long Context Not comparable

Codellama 70b Instruct: —, Mixtral 8x7B: 33.4 (#260)

Long Context benchmarks
BenchmarkCodellama 70b InstructMixtral 8x7B
LMArena Longer Query—1103

Writing & Preference Too close to call

Codellama 70b Instruct: 33.4 (#277), Mixtral 8x7B: 34.2 (#270)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructMixtral 8x7B
LMArena Text10571132
LMArena Creative Writing—1109
WildBench—67.3%
LMArena Multi-Turn—1115

Frequently asked questions

Is Codellama 70b Instruct better than Mixtral 8x7B?

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 27.1 on the Noometry Index.

Is Codellama 70b Instruct or Mixtral 8x7B better for coding?

Codellama 70b Instruct scores higher on coding benchmarks: 37.6 versus 32.8 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Mixtral 8x7B share?

5 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Mixtral 8x7B has 38.

Related comparisons

Go deeper