Model comparison

Kimi K2 Thinking Turbo vs Llama-3.3-70B-Instruct

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 30.6 on the Noometry Index.

Last verified . 19 shared benchmarks.

Kimi K2 Thinking Turbo Moonshot AI

45.8

Rank #70 Confirmed

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Kimi K2 Thinking Turbo scores higher in 8 categories and Llama-3.3-70B-Instruct in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2 Thinking Turbo leads 47.4 to 15.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 83.1% for Kimi K2 Thinking Turbo and 5.1% for Llama-3.3-70B-Instruct.

Side by side

Kimi K2 Thinking Turbo and Llama-3.3-70B-Instruct specifications
Kimi K2 Thinking TurboLlama-3.3-70B-Instruct
ProviderMoonshot AIMeta
Noometry Index45.830.6
Released2025-11-062024-12-06
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.32
Results tracked2143

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 38.4 (#178), Llama-3.3-70B-Instruct: 31.0 (#290)

Coding benchmarks
BenchmarkKimi K2 Thinking TurboLlama-3.3-70B-Instruct
LMArena Coding14541268
LMArena WebDev1322—
SciCode—26%
WeirdML—14.4%
BigCodeBench Instruct—46.9%
LiveBench Coding—36.6%
BigCodeBench Complete—57.5%

Agentic & Tool Use Not comparable

Kimi K2 Thinking Turbo: —, Llama-3.3-70B-Instruct: 25.8 (#105)

Agentic & Tool Use benchmarks
BenchmarkKimi K2 Thinking TurboLlama-3.3-70B-Instruct
Berkeley Function Calling Leaderboard—31.9%
BALROG—23%

Reasoning Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 32.5 (#75), Llama-3.3-70B-Instruct: 14.1 (#327)

Reasoning benchmarks
BenchmarkKimi K2 Thinking TurboLlama-3.3-70B-Instruct
LMArena Hard Prompts14281257
SimpleBench—19.9%
CritPt—0%
Chess Puzzles20%—
LiveBench Reasoning—50.8%
DTBench—59.5%
LiveBench Data Analysis—49.5%
LMCA—17.5%
Epoch Capabilities Index—127.33
ForecastBench—58.6
LiveBench—50.2%

Math Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 47.4 (#68), Llama-3.3-70B-Instruct: 15.3 (#298)

Math benchmarks
BenchmarkKimi K2 Thinking TurboLlama-3.3-70B-Instruct
OTIS Mock AIME 2024-202583.1%5.1%
LMArena Math14291267
LiveBench Math—42.2%
MATH Level 5—41.6%

Knowledge Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 50.9 (#69), Llama-3.3-70B-Instruct: 30.6 (#226)

Knowledge benchmarks
BenchmarkKimi K2 Thinking TurboLlama-3.3-70B-Instruct
GPQA Diamond84.2%47.4%
LMArena Expert14391225
Confabulations—22.8%
Vectara Hallucination Rate—4.1%
MMLU—86.3%

Multilingual Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 51.4 (#109), Llama-3.3-70B-Instruct: 39.9 (#220)

Multilingual benchmarks
BenchmarkKimi K2 Thinking TurboLlama-3.3-70B-Instruct
LMArena Non-English13981236
LMArena Chinese14561217
LMArena French14251281
LMArena German13901251
LMArena Japanese13571150
LMArena Korean13301143
LMArena Russian13911252
LMArena Spanish14061270

Instruction Following Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 74.0 (#109), Llama-3.3-70B-Instruct: 71.1 (#157)

Instruction Following benchmarks
BenchmarkKimi K2 Thinking TurboLlama-3.3-70B-Instruct
LMArena Instruction Following14031242
LiveBench Instruction Following—82.7%

Long Context Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 43.2 (#102), Llama-3.3-70B-Instruct: 26.4 (#295)

Long Context benchmarks
BenchmarkKimi K2 Thinking TurboLlama-3.3-70B-Instruct
LMArena Longer Query14151256
Fiction.LiveBench—33.3%

Writing & Preference Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 60.0 (#104), Llama-3.3-70B-Instruct: 47.6 (#207)

Writing & Preference benchmarks
BenchmarkKimi K2 Thinking TurboLlama-3.3-70B-Instruct
LMArena Text14151274
LMArena Creative Writing13741250
LMArena Multi-Turn14141280
LiveBench Language—39.2%

Frequently asked questions

Is Kimi K2 Thinking Turbo better than Llama-3.3-70B-Instruct?

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 30.6 on the Noometry Index.

Is Kimi K2 Thinking Turbo or Llama-3.3-70B-Instruct better for coding?

Kimi K2 Thinking Turbo scores higher on coding benchmarks: 38.4 versus 31.0 in the Noometry coding category.

How many benchmarks do Kimi K2 Thinking Turbo and Llama-3.3-70B-Instruct share?

19 benchmarks have published results for both models. Kimi K2 Thinking Turbo has 21 scored results on Noometry and Llama-3.3-70B-Instruct has 43.

Related comparisons

Go deeper