Model comparison

Kimi K2 Thinking Turbo vs Mistral Small 3.2

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 31.2 on the Noometry Index.

Last verified . 3 shared benchmarks.

Kimi K2 Thinking Turbo Moonshot AI

45.8

Rank #70 Confirmed

Mistral Small 3.2 Mistral AI

31.2

Rank #280 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Kimi K2 Thinking Turbo scores higher in 4 categories and Mistral Small 3.2 in 0 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Kimi K2 Thinking Turbo leads 50.9 to 26.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 83.1% for Kimi K2 Thinking Turbo and 30.3% for Mistral Small 3.2.

Side by side

Kimi K2 Thinking Turbo and Mistral Small 3.2 specifications
Kimi K2 Thinking TurboMistral Small 3.2
ProviderMoonshot AIMistral AI
Noometry Index45.831.2
Released2025-11-062025-06-20
WeightsOpenOpen
Context window—256K
Max output—16K
Input $ / M tokens—$0.0938
Output $ / M tokens—$0.25
Results tracked216

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Kimi K2 Thinking Turbo: 38.4 (#178), Mistral Small 3.2: —

Coding benchmarks
BenchmarkKimi K2 Thinking TurboMistral Small 3.2
LMArena WebDev1322—
LMArena Coding1454—

Reasoning Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 32.5 (#75), Mistral Small 3.2: 18.1 (#287)

Reasoning benchmarks
BenchmarkKimi K2 Thinking TurboMistral Small 3.2
Chess Puzzles20%1%
Kagi LLM Benchmark—40.4%
LMArena Hard Prompts1428—
Epoch Capabilities Index—131.74

Math Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 47.4 (#68), Mistral Small 3.2: 26.3 (#260)

Math benchmarks
BenchmarkKimi K2 Thinking TurboMistral Small 3.2
OTIS Mock AIME 2024-202583.1%30.3%
LMArena Math1429—

Knowledge Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 50.9 (#69), Mistral Small 3.2: 26.7 (#256)

Knowledge benchmarks
BenchmarkKimi K2 Thinking TurboMistral Small 3.2
GPQA Diamond84.2%49.1%
LMArena Expert1439—

Multilingual Not comparable

Kimi K2 Thinking Turbo: 51.4 (#109), Mistral Small 3.2: —

Multilingual benchmarks
BenchmarkKimi K2 Thinking TurboMistral Small 3.2
LMArena Non-English1398—
LMArena Chinese1456—
LMArena French1425—
LMArena German1390—
LMArena Japanese1357—
LMArena Korean1330—
LMArena Russian1391—
LMArena Spanish1406—

Instruction Following Not comparable

Kimi K2 Thinking Turbo: 74.0 (#109), Mistral Small 3.2: —

Instruction Following benchmarks
BenchmarkKimi K2 Thinking TurboMistral Small 3.2
LMArena Instruction Following1403—

Long Context Not comparable

Kimi K2 Thinking Turbo: 43.2 (#102), Mistral Small 3.2: —

Long Context benchmarks
BenchmarkKimi K2 Thinking TurboMistral Small 3.2
LMArena Longer Query1415—

Writing & Preference Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 60.0 (#104), Mistral Small 3.2: 45.0 (#224)

Writing & Preference benchmarks
BenchmarkKimi K2 Thinking TurboMistral Small 3.2
LMArena Text1415—
LMArena Creative Writing1374—
EQ-Bench Creative Writing—1255
LMArena Multi-Turn1414—

Frequently asked questions

Is Kimi K2 Thinking Turbo better than Mistral Small 3.2?

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 31.2 on the Noometry Index.

How many benchmarks do Kimi K2 Thinking Turbo and Mistral Small 3.2 share?

3 benchmarks have published results for both models. Kimi K2 Thinking Turbo has 21 scored results on Noometry and Mistral Small 3.2 has 6.

Related comparisons

Go deeper