Model comparison

Kimi K2 Thinking Turbo vs Mistral Large 4

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 43.1 on the Noometry Index.

Last verified . 13 shared benchmarks.

Kimi K2 Thinking Turbo Moonshot AI

45.8

Rank #70 Confirmed

Mistral Large 4 Mistral AI

43.1

Rank #99 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Kimi K2 Thinking Turbo scores higher in 3 categories and Mistral Large 4 in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Kimi K2 Thinking Turbo leads 50.9 to 36.6.
  • Kimi K2 Thinking Turbo has downloadable open weights; the other is API-only.

Side by side

Kimi K2 Thinking Turbo and Mistral Large 4 specifications
Kimi K2 Thinking TurboMistral Large 4
ProviderMoonshot AIMistral AI
Noometry Index45.843.1
Released2025-11-062026-10-06
WeightsOpenProprietary
Context window—1.05M
Max output—262K
Input $ / M tokens—$0.68
Output $ / M tokens—$2.09
Results tracked2115

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large 4 leads

Kimi K2 Thinking Turbo: 38.4 (#178), Mistral Large 4: 48.6 (#57)

Coding benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large 4
LMArena WebDev13221541
LMArena Coding14541475

Reasoning Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 32.5 (#75), Mistral Large 4: 22.5 (#192)

Reasoning benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large 4
LMArena Hard Prompts14281444
NYT Connections (extended)—27.4%
Chess Puzzles20%—

Math Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 47.4 (#68), Mistral Large 4: 40.4 (#91)

Math benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large 4
LMArena Math14291488
OTIS Mock AIME 2024-202583.1%—

Knowledge Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 50.9 (#69), Mistral Large 4: 36.6 (#166)

Knowledge benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large 4
LMArena Expert14391447
GPQA Diamond84.2%—
SimpleQA Verified—20%

Multilingual Mistral Large 4 leads

Kimi K2 Thinking Turbo: 51.4 (#109), Mistral Large 4: 52.6 (#82)

Multilingual benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large 4
LMArena Non-English13981415
LMArena Chinese14561491
LMArena Russian13911414
LMArena French1425—
LMArena German1390—
LMArena Japanese1357—
LMArena Korean1330—
LMArena Spanish1406—

Instruction Following Mistral Large 4 leads

Kimi K2 Thinking Turbo: 74.0 (#109), Mistral Large 4: 75.0 (#76)

Instruction Following benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large 4
LMArena Instruction Following14031424

Long Context Too close to call

Kimi K2 Thinking Turbo: 43.2 (#102), Mistral Large 4: 43.6 (#89)

Long Context benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large 4
LMArena Longer Query14151429

Writing & Preference Too close to call

Kimi K2 Thinking Turbo: 60.0 (#104), Mistral Large 4: 60.4 (#97)

Writing & Preference benchmarks
BenchmarkKimi K2 Thinking TurboMistral Large 4
LMArena Text14151427
LMArena Creative Writing13741361
LMArena Multi-Turn14141424

Frequently asked questions

Is Kimi K2 Thinking Turbo better than Mistral Large 4?

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 43.1 on the Noometry Index.

Is Kimi K2 Thinking Turbo or Mistral Large 4 better for coding?

Mistral Large 4 scores higher on coding benchmarks: 48.6 versus 38.4 in the Noometry coding category.

How many benchmarks do Kimi K2 Thinking Turbo and Mistral Large 4 share?

13 benchmarks have published results for both models. Kimi K2 Thinking Turbo has 21 scored results on Noometry and Mistral Large 4 has 15.

Related comparisons

Go deeper