Model comparison

Kimi K2 Thinking Turbo vs Mistral Medium

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 36.3 on the Noometry Index.

Last verified . 19 shared benchmarks.

Kimi K2 Thinking Turbo Moonshot AI

45.8

Rank #70 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Kimi K2 Thinking Turbo scores higher in 6 categories and Mistral Medium in 2 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Kimi K2 Thinking Turbo leads 50.9 to 25.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 83.1% for Kimi K2 Thinking Turbo and 32.2% for Mistral Medium.

Side by side

Kimi K2 Thinking Turbo and Mistral Medium specifications
Kimi K2 Thinking TurboMistral Medium
ProviderMoonshot AIMistral AI
Noometry Index45.836.3
Released2025-11-062023-12-11
WeightsOpenOpen
Context window—262K
Max output—262K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked2136

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 38.4 (#178), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkKimi K2 Thinking TurboMistral Medium
LMArena Coding14541434
FrontierCode—8%
LMArena WebDev1322—
SciCode—40.2%
WeirdML—43.7%
ALE-Bench—763.98

Agentic & Tool Use Not comparable

Kimi K2 Thinking Turbo: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkKimi K2 Thinking TurboMistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 32.5 (#75), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkKimi K2 Thinking TurboMistral Medium
LMArena Hard Prompts14281426
Kagi LLM Benchmark—50%
CritPt—0%
Chess Puzzles20%—
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 47.4 (#68), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkKimi K2 Thinking TurboMistral Medium
OTIS Mock AIME 2024-202583.1%32.2%
LMArena Math14291408
ProofBench—9%
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 50.9 (#69), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkKimi K2 Thinking TurboMistral Medium
GPQA Diamond84.2%59.5%
LMArena Expert14391408
Humanity's Last Exam—4.5%
Vectara Hallucination Rate—22.7%

Multimodal Not comparable

Kimi K2 Thinking Turbo: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkKimi K2 Thinking TurboMistral Medium
LMArena Vision—1172

Multilingual Too close to call

Kimi K2 Thinking Turbo: 51.4 (#109), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkKimi K2 Thinking TurboMistral Medium
LMArena Non-English13981408
LMArena Chinese14561447
LMArena French14251459
LMArena German13901432
LMArena Japanese13571378
LMArena Korean13301380
LMArena Russian13911411
LMArena Spanish14061433

Instruction Following Too close to call

Kimi K2 Thinking Turbo: 74.0 (#109), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkKimi K2 Thinking TurboMistral Medium
LMArena Instruction Following14031398

Long Context Too close to call

Kimi K2 Thinking Turbo: 43.2 (#102), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkKimi K2 Thinking TurboMistral Medium
LMArena Longer Query14151406

Writing & Preference Too close to call

Kimi K2 Thinking Turbo: 60.0 (#104), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkKimi K2 Thinking TurboMistral Medium
LMArena Text14151424
LMArena Creative Writing13741391
LMArena Multi-Turn14141418
Short-Story Creative Writing—77.3%

Frequently asked questions

Is Kimi K2 Thinking Turbo better than Mistral Medium?

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 36.3 on the Noometry Index.

Is Kimi K2 Thinking Turbo or Mistral Medium better for coding?

Kimi K2 Thinking Turbo scores higher on coding benchmarks: 38.4 versus 34.2 in the Noometry coding category.

How many benchmarks do Kimi K2 Thinking Turbo and Mistral Medium share?

19 benchmarks have published results for both models. Kimi K2 Thinking Turbo has 21 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper