Model comparison

Kimi K2.7 Code vs Mixtral 8x22B

Kimi K2.7 Code is the stronger model overall, scoring 43.3 to 27.1 on the Noometry Index.

Last verified . 3 shared benchmarks.

Kimi K2.7 Code Moonshot AI

43.3

Rank #94 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Kimi K2.7 Code scores higher in 5 categories and Mixtral 8x22B in 0 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Kimi K2.7 Code leads 53.5 to 15.1.
  • The biggest single-benchmark swing is GPQA Diamond: 87.9% for Kimi K2.7 Code and 34.1% for Mixtral 8x22B.
  • Kimi K2.7 Code is cheaper at $0.95 / $4 per million input/output tokens, against $2 / $6 for Mixtral 8x22B.
  • Kimi K2.7 Code accepts more context: 262K tokens versus 64K.

Side by side

Kimi K2.7 Code and Mixtral 8x22B specifications
Kimi K2.7 CodeMixtral 8x22B
ProviderMoonshot AIMistral AI
Noometry Index43.327.1
Released2026-06-122024-04-17
WeightsOpenOpen
Context window262K64K
Max output262K64K
Input $ / M tokens$0.95$2
Output $ / M tokens$4$6
Results tracked1934

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.7 Code leads

Kimi K2.7 Code: 42.9 (#95), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkKimi K2.7 CodeMixtral 8x22B
WeirdML54.1%3.2%
DeepSWE30.5%—
FrontierCode30.1%—
LMArena WebDev1473—
SciCode47.5%—
BigCodeBench Instruct—40.6%
LMArena Coding—1166
BigCodeBench Complete—50.2%
ALE-Bench886.23—
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Too close to call

Kimi K2.7 Code: 24.0 (#122), Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkKimi K2.7 CodeMixtral 8x22B
APEX-Agents37.6%—
Cybench—7.5%
GBAEval0.9%—
Vending-Bench 25,083—

Reasoning Kimi K2.7 Code leads

Kimi K2.7 Code: 39.0 (#61), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkKimi K2.7 CodeMixtral 8x22B
Epoch Capabilities Index149.97122.03
SimpleBench57.9%—
CritPt10%—
Chess Puzzles21%—
LMArena Hard Prompts—1150
DTBench—55.1%
Surface Evolver Bench48.8%—
ForecastBench—56.3

Math Kimi K2.7 Code leads

Kimi K2.7 Code: 52.9 (#48), Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkKimi K2.7 CodeMixtral 8x22B
FrontierMath (Tiers 1-3)54%—
FrontierMath Tier 412.2%—
OTIS Mock AIME 2024-202595.6%—
Omni-MATH—16.3%
LMArena Math—1184
MATH Level 5—24.2%

Knowledge Kimi K2.7 Code leads

Kimi K2.7 Code: 53.5 (#57), Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkKimi K2.7 CodeMixtral 8x22B
GPQA Diamond87.9%34.1%
SimpleQA Verified36.5%—
MMLU-Pro—46%
GPQA (HELM)—33.4%
LMArena Expert—1113
MMLU—77.8%

Multilingual Not comparable

Kimi K2.7 Code: —, Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkKimi K2.7 CodeMixtral 8x22B
LMArena Non-English—1128
LMArena Chinese—1116
LMArena French—1166
LMArena German—1141
LMArena Japanese—1037
LMArena Korean—1057
LMArena Russian—1158
LMArena Spanish—1151

Instruction Following Not comparable

Kimi K2.7 Code: —, Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkKimi K2.7 CodeMixtral 8x22B
IFEval—72.4%
LMArena Instruction Following—1147

Long Context Not comparable

Kimi K2.7 Code: —, Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkKimi K2.7 CodeMixtral 8x22B
LMArena Longer Query—1144

Writing & Preference Not comparable

Kimi K2.7 Code: —, Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkKimi K2.7 CodeMixtral 8x22B
LMArena Text—1162
LMArena Creative Writing—1141
WildBench—71.1%
LMArena Multi-Turn—1130

Frequently asked questions

Is Kimi K2.7 Code better than Mixtral 8x22B?

Kimi K2.7 Code is the stronger model overall, scoring 43.3 to 27.1 on the Noometry Index.

Which is cheaper, Kimi K2.7 Code or Mixtral 8x22B?

Kimi K2.7 Code is cheaper. It lists at $0.95 per million input tokens and $4 per million output tokens; Mixtral 8x22B lists at $2 and $6.

Is Kimi K2.7 Code or Mixtral 8x22B better for coding?

Kimi K2.7 Code scores higher on coding benchmarks: 42.9 versus 24.2 in the Noometry coding category.

Which has the bigger context window?

Kimi K2.7 Code does, with 262K tokens against 64K.

How many benchmarks do Kimi K2.7 Code and Mixtral 8x22B share?

3 benchmarks have published results for both models. Kimi K2.7 Code has 19 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper