Model comparison

Kimi K2.6 vs Mistral Large

Kimi K2.6 is the stronger model overall, scoring 47.7 to 31.9 on the Noometry Index.

Last verified . 28 shared benchmarks.

Kimi K2.6 Moonshot AI

47.7

Rank #60 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Kimi K2.6 scores higher in 8 categories and Mistral Large in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2.6 leads 57.0 to 18.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 96.1% for Kimi K2.6 and 8.5% for Mistral Large.
  • Kimi K2.6 is cheaper at $0.95 / $4 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Kimi K2.6 accepts more context: 262K tokens versus 131K.

Side by side

Kimi K2.6 and Mistral Large specifications
Kimi K2.6Mistral Large
ProviderMoonshot AIMistral AI
Noometry Index47.731.9
Released2026-04-202024-02-26
WeightsOpenOpen
Context window262K131K
Max output262K16K
Input $ / M tokens$0.95$2
Output $ / M tokens$4$6
Results tracked5151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.6 leads

Kimi K2.6: 50.7 (#43), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkKimi K2.6Mistral Large
SciCode53.5%36.2%
LMArena Coding14881277
ALE-Bench1,093264.7
SWE-bench Verified76.7%—
LMArena WebDev1509—
WeirdML55.9%—
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Mistral Large leads

Kimi K2.6: 21.9 (#137), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkKimi K2.6Mistral Large
Berkeley Function Calling Leaderboard—38.4%
OSWorld 2.04.6%—
ExploitBench18.4%—
GBAEval0.9%—
GDP.pdf12%—
Vending-Bench 26,205—

Reasoning Kimi K2.6 leads

Kimi K2.6: 40.5 (#55), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkKimi K2.6Mistral Large
CritPt8%0%
LMArena Hard Prompts14701257
DTBench90.9%65.1%
LMCA37.3%16.7%
Epoch Capabilities Index151.05128.52
SimpleBench—22.5%
NYT Connections (extended)87.2%—
Chess Puzzles26%—
EBR-Bench2.4%—
LiveBench Reasoning—43.5%
Mystery Game Puzzles18%—
LiveBench Data Analysis—50.1%
ForecastBench—57.1
LiveBench—48.4%

Math Kimi K2.6 leads

Kimi K2.6: 57.0 (#41), Mistral Large: 18.2 (#291)

Knowledge Kimi K2.6 leads

Kimi K2.6: 54.0 (#54), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkKimi K2.6Mistral Large
GPQA Diamond90.8%51.3%
Vectara Hallucination Rate10.8%4.5%
LMArena Expert14911232
SimpleQA Verified34.9%—
MMLU-Pro—59.9%
Confabulations—21.4%
GPQA (HELM)—43.5%
MMLU—80%

Multimodal Not comparable

Kimi K2.6: 31.6 (#103), Mistral Large: —

Multimodal benchmarks
BenchmarkKimi K2.6Mistral Large
LMArena Vision1283—
Blueprint-Bench 23.9%—
Furniture Assembly21.7%—
LMArena Document1451—

Multilingual Kimi K2.6 leads

Kimi K2.6: 54.9 (#37), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkKimi K2.6Mistral Large
LMArena Non-English14461237
LMArena Chinese15211240
LMArena French14711325
LMArena German14501254
LMArena Japanese14431188
LMArena Korean14271202
LMArena Russian14461257
LMArena Spanish14641268

Instruction Following Kimi K2.6 leads

Kimi K2.6: 76.3 (#43), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkKimi K2.6Mistral Large
LMArena Instruction Following14511249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Kimi K2.6 leads

Kimi K2.6: 44.9 (#52), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkKimi K2.6Mistral Large
LMArena Longer Query14681261

Writing & Preference Kimi K2.6 leads

Kimi K2.6: 68.5 (#26), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkKimi K2.6Mistral Large
LMArena Text14551266
LMArena Creative Writing14341243
EQ-Bench Creative Writing1725985
LMArena Multi-Turn14531260
Short-Story Creative Writing—69%
WildBench—80.1%
EQ-Bench 41202—
LiveBench Language—39.4%

Frequently asked questions

Is Kimi K2.6 better than Mistral Large?

Kimi K2.6 is the stronger model overall, scoring 47.7 to 31.9 on the Noometry Index.

Which is cheaper, Kimi K2.6 or Mistral Large?

Kimi K2.6 is cheaper. It lists at $0.95 per million input tokens and $4 per million output tokens; Mistral Large lists at $2 and $6.

Is Kimi K2.6 or Mistral Large better for coding?

Kimi K2.6 scores higher on coding benchmarks: 50.7 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Kimi K2.6 does, with 262K tokens against 131K.

How many benchmarks do Kimi K2.6 and Mistral Large share?

28 benchmarks have published results for both models. Kimi K2.6 has 51 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper