Model comparison

Codestral vs Mixtral 8x22B

Codestral is the stronger model overall, scoring 30.6 to 27.1 on the Noometry Index.

Last verified . 4 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Codestral scores higher in 1 category and Mixtral 8x22B in 1 category; one gap is clear of the uncertainty.
  • The widest gap is in coding, where Codestral leads 27.3 to 24.2.
  • Codestral is cheaper at $0.30 / $0.90 per million input/output tokens, against $2 / $6 for Mixtral 8x22B.
  • Codestral accepts more context: 256K tokens versus 64K.
  • Mixtral 8x22B has downloadable open weights; the other is API-only.

Side by side

Codestral and Mixtral 8x22B specifications
CodestralMixtral 8x22B
ProviderMistral AIMistral AI
Noometry Index30.627.1
Released2024-05-292024-04-17
WeightsProprietaryOpen
Context window256K64K
Max output8K64K
Input $ / M tokens$0.30$2
Output $ / M tokens$0.90$6
Results tracked734

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codestral leads

Codestral: 27.3 (#321), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkCodestralMixtral 8x22B
BigCodeBench Instruct41.8%40.6%
BigCodeBench Complete52.5%50.2%
HumanEval+73.8%72%
MBPP+61.9%64.3%
Aider Polyglot11.1%—
WeirdML—3.2%
LMArena Coding—1166
ALE-Bench137.78—

Agentic & Tool Use Not comparable

Codestral: —, Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkCodestralMixtral 8x22B
Cybench—7.5%

Reasoning Too close to call

Codestral: 19.8 (#251), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkCodestralMixtral 8x22B
Kagi LLM Benchmark32.5%—
LMArena Hard Prompts—1150
DTBench—55.1%
Epoch Capabilities Index—122.03
ForecastBench—56.3

Math Not comparable

Codestral: —, Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkCodestralMixtral 8x22B
Omni-MATH—16.3%
LMArena Math—1184
MATH Level 5—24.2%

Knowledge Not comparable

Codestral: —, Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkCodestralMixtral 8x22B
GPQA Diamond—34.1%
MMLU-Pro—46%
GPQA (HELM)—33.4%
LMArena Expert—1113
MMLU—77.8%

Multilingual Not comparable

Codestral: —, Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkCodestralMixtral 8x22B
LMArena Non-English—1128
LMArena Chinese—1116
LMArena French—1166
LMArena German—1141
LMArena Japanese—1037
LMArena Korean—1057
LMArena Russian—1158
LMArena Spanish—1151

Instruction Following Not comparable

Codestral: —, Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkCodestralMixtral 8x22B
IFEval—72.4%
LMArena Instruction Following—1147

Long Context Not comparable

Codestral: —, Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkCodestralMixtral 8x22B
LMArena Longer Query—1144

Writing & Preference Not comparable

Codestral: —, Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkCodestralMixtral 8x22B
LMArena Text—1162
LMArena Creative Writing—1141
WildBench—71.1%
LMArena Multi-Turn—1130

Frequently asked questions

Is Codestral better than Mixtral 8x22B?

Codestral is the stronger model overall, scoring 30.6 to 27.1 on the Noometry Index.

Which is cheaper, Codestral or Mixtral 8x22B?

Codestral is cheaper. It lists at $0.30 per million input tokens and $0.90 per million output tokens; Mixtral 8x22B lists at $2 and $6.

Is Codestral or Mixtral 8x22B better for coding?

Codestral scores higher on coding benchmarks: 27.3 versus 24.2 in the Noometry coding category.

Which has the bigger context window?

Codestral does, with 256K tokens against 64K.

How many benchmarks do Codestral and Mixtral 8x22B share?

4 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper