Model comparison

Codestral vs Kimi K2 Thinking Turbo

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 30.6 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Kimi K2 Thinking Turbo Moonshot AI

45.8

Rank #70 Confirmed

Summary

  • The widest gap is in reasoning, where Kimi K2 Thinking Turbo leads 32.5 to 19.8.
  • Kimi K2 Thinking Turbo has downloadable open weights; the other is API-only.

Side by side

Codestral and Kimi K2 Thinking Turbo specifications
CodestralKimi K2 Thinking Turbo
ProviderMistral AIMoonshot AI
Noometry Index30.645.8
Released2024-05-292025-11-06
WeightsProprietaryOpen
Context window256K—
Max output8K—
Input $ / M tokens$0.30—
Output $ / M tokens$0.90—
Results tracked721

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 Thinking Turbo leads

Codestral: 27.3 (#321), Kimi K2 Thinking Turbo: 38.4 (#178)

Coding benchmarks
BenchmarkCodestralKimi K2 Thinking Turbo
Aider Polyglot11.1%—
LMArena WebDev—1322
BigCodeBench Instruct41.8%—
LMArena Coding—1454
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Reasoning Kimi K2 Thinking Turbo leads

Codestral: 19.8 (#251), Kimi K2 Thinking Turbo: 32.5 (#75)

Reasoning benchmarks
BenchmarkCodestralKimi K2 Thinking Turbo
Kagi LLM Benchmark32.5%—
Chess Puzzles—20%
LMArena Hard Prompts—1428

Math Not comparable

Codestral: —, Kimi K2 Thinking Turbo: 47.4 (#68)

Math benchmarks
BenchmarkCodestralKimi K2 Thinking Turbo
OTIS Mock AIME 2024-2025—83.1%
LMArena Math—1429

Knowledge Not comparable

Codestral: —, Kimi K2 Thinking Turbo: 50.9 (#69)

Knowledge benchmarks
BenchmarkCodestralKimi K2 Thinking Turbo
GPQA Diamond—84.2%
LMArena Expert—1439

Multilingual Not comparable

Codestral: —, Kimi K2 Thinking Turbo: 51.4 (#109)

Multilingual benchmarks
BenchmarkCodestralKimi K2 Thinking Turbo
LMArena Non-English—1398
LMArena Chinese—1456
LMArena French—1425
LMArena German—1390
LMArena Japanese—1357
LMArena Korean—1330
LMArena Russian—1391
LMArena Spanish—1406

Instruction Following Not comparable

Codestral: —, Kimi K2 Thinking Turbo: 74.0 (#109)

Instruction Following benchmarks
BenchmarkCodestralKimi K2 Thinking Turbo
LMArena Instruction Following—1403

Long Context Not comparable

Codestral: —, Kimi K2 Thinking Turbo: 43.2 (#102)

Long Context benchmarks
BenchmarkCodestralKimi K2 Thinking Turbo
LMArena Longer Query—1415

Writing & Preference Not comparable

Codestral: —, Kimi K2 Thinking Turbo: 60.0 (#104)

Writing & Preference benchmarks
BenchmarkCodestralKimi K2 Thinking Turbo
LMArena Text—1415
LMArena Creative Writing—1374
LMArena Multi-Turn—1414

Frequently asked questions

Is Codestral better than Kimi K2 Thinking Turbo?

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 30.6 on the Noometry Index.

Is Codestral or Kimi K2 Thinking Turbo better for coding?

Kimi K2 Thinking Turbo scores higher on coding benchmarks: 38.4 versus 27.3 in the Noometry coding category.

How many benchmarks do Codestral and Kimi K2 Thinking Turbo share?

0 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Kimi K2 Thinking Turbo has 21.

Related comparisons

Go deeper