Model comparison

Codestral vs Olmo 3 32b Think

Olmo 3 32b Think is the stronger model overall, scoring 38.7 to 30.6 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Summary

  • The widest gap is in coding, where Olmo 3 32b Think leads 38.6 to 27.3.
  • Olmo 3 32b Think has downloadable open weights; the other is API-only.

Side by side

Codestral and Olmo 3 32b Think specifications
CodestralOlmo 3 32b Think
ProviderMistral AIAllen Institute for AI (Ai2)
Noometry Index30.638.7
Released2024-05-29—
WeightsProprietaryOpen
Context window256K—
Max output8K—
Input $ / M tokens$0.30—
Output $ / M tokens$0.90—
Results tracked714

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3 32b Think leads

Codestral: 27.3 (#321), Olmo 3 32b Think: 38.6 (#172)

Coding benchmarks
BenchmarkCodestralOlmo 3 32b Think
Aider Polyglot11.1%—
BigCodeBench Instruct41.8%—
LMArena Coding—1319
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Reasoning Olmo 3 32b Think leads

Codestral: 19.8 (#251), Olmo 3 32b Think: 25.9 (#140)

Reasoning benchmarks
BenchmarkCodestralOlmo 3 32b Think
Kagi LLM Benchmark32.5%—
LMArena Hard Prompts—1302

Math Not comparable

Codestral: —, Olmo 3 32b Think: 36.5 (#165)

Math benchmarks
BenchmarkCodestralOlmo 3 32b Think
LMArena Math—1316

Knowledge Not comparable

Codestral: —, Olmo 3 32b Think: 35.0 (#190)

Knowledge benchmarks
BenchmarkCodestralOlmo 3 32b Think
LMArena Expert—1273

Multilingual Not comparable

Codestral: —, Olmo 3 32b Think: 41.2 (#210)

Multilingual benchmarks
BenchmarkCodestralOlmo 3 32b Think
LMArena Non-English—1255
LMArena Chinese—1300
LMArena French—1291
LMArena German—1290
LMArena Russian—1254

Instruction Following Not comparable

Codestral: —, Olmo 3 32b Think: 67.2 (#198)

Instruction Following benchmarks
BenchmarkCodestralOlmo 3 32b Think
LMArena Instruction Following—1275

Long Context Not comparable

Codestral: —, Olmo 3 32b Think: 39.4 (#182)

Long Context benchmarks
BenchmarkCodestralOlmo 3 32b Think
LMArena Longer Query—1296

Writing & Preference Not comparable

Codestral: —, Olmo 3 32b Think: 49.1 (#193)

Writing & Preference benchmarks
BenchmarkCodestralOlmo 3 32b Think
LMArena Text—1300
LMArena Creative Writing—1256
LMArena Multi-Turn—1290

Frequently asked questions

Is Codestral better than Olmo 3 32b Think?

Olmo 3 32b Think is the stronger model overall, scoring 38.7 to 30.6 on the Noometry Index.

Is Codestral or Olmo 3 32b Think better for coding?

Olmo 3 32b Think scores higher on coding benchmarks: 38.6 versus 27.3 in the Noometry coding category.

How many benchmarks do Codestral and Olmo 3 32b Think share?

0 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Olmo 3 32b Think has 14.

Related comparisons

Go deeper