Model comparison

Kimi K2 Thinking Turbo vs Mistral Large

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 31.9 on the Noometry Index.

Last verified . 19 shared benchmarks.

Kimi K2 Thinking Turbo Moonshot AI

45.8

Rank #70 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Kimi K2 Thinking Turbo scores higher in 8 categories and Mistral Large in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2 Thinking Turbo leads 47.4 to 18.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 83.1% for Kimi K2 Thinking Turbo and 8.5% for Mistral Large.

Side by side

Kimi K2 Thinking Turbo and Mistral Large specifications
Kimi K2 Thinking TurboMistral Large
ProviderMoonshot AIMistral AI
Noometry Index45.831.9
Released2025-11-062024-02-26
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked2151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 38.4 (#178), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large
LMArena Coding14541277
LMArena WebDev1322—
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Kimi K2 Thinking Turbo: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 32.5 (#75), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large
LMArena Hard Prompts14281257
SimpleBench—22.5%
CritPt—0%
Chess Puzzles20%—
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 47.4 (#68), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large
OTIS Mock AIME 2024-202583.1%8.5%
LMArena Math14291262
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 50.9 (#69), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large
GPQA Diamond84.2%51.3%
LMArena Expert14391232
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
MMLU—80%

Multilingual Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 51.4 (#109), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large
LMArena Non-English13981237
LMArena Chinese14561240
LMArena French14251325
LMArena German13901254
LMArena Japanese13571188
LMArena Korean13301202
LMArena Russian13911257
LMArena Spanish14061268

Instruction Following Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 74.0 (#109), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large
LMArena Instruction Following14031249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 43.2 (#102), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large
LMArena Longer Query14151261

Writing & Preference Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 60.0 (#104), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large
LMArena Text14151266
LMArena Creative Writing13741243
LMArena Multi-Turn14141260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Kimi K2 Thinking Turbo better than Mistral Large?

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 31.9 on the Noometry Index.

Is Kimi K2 Thinking Turbo or Mistral Large better for coding?

Kimi K2 Thinking Turbo scores higher on coding benchmarks: 38.4 versus 34.3 in the Noometry coding category.

How many benchmarks do Kimi K2 Thinking Turbo and Mistral Large share?

19 benchmarks have published results for both models. Kimi K2 Thinking Turbo has 21 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper