Model comparison

gpt-oss-120b vs Kimi K2 (Jul 2025)

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 36.3 on the Noometry Index. gpt-oss-120b costs 14× less per token, which makes it the better buy when Kimi K2 (Jul 2025)'s lead doesn't matter for your workload.

Last verified . 36 shared benchmarks.

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 36 benchmarks with published results for both. gpt-oss-120b scores higher in 2 categories and Kimi K2 (Jul 2025) in 7 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Kimi K2 (Jul 2025) leads 32.4 to 12.2.
  • The biggest single-benchmark swing is SWE-bench Verified (bash only): 26% for gpt-oss-120b and 63.4% for Kimi K2 (Jul 2025).
  • gpt-oss-120b is cheaper at $0.037 / $0.17 per million input/output tokens, against $0.57 / $2.30 for Kimi K2 (Jul 2025).
  • Kimi K2 (Jul 2025) accepts more context: 262K tokens versus 131K.

Side by side

gpt-oss-120b and Kimi K2 (Jul 2025) specifications
gpt-oss-120bKimi K2 (Jul 2025)
ProviderOpenAIMoonshot AI
Noometry Index36.341.2
Released2025-08-052025-07-12
WeightsOpenOpen
Context window131K262K
Max output41K262K
Input $ / M tokens$0.037$0.57
Output $ / M tokens$0.17$2.30
Results tracked4842

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 (Jul 2025) leads

gpt-oss-120b: 33.5 (#256), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
Benchmarkgpt-oss-120bKimi K2 (Jul 2025)
SWE-bench Verified (bash only)26%63.4%
Aider Polyglot41.8%59.1%
WeirdML48.2%42.8%
LMArena Coding13801399
ALE-Bench575.62597.5
SciCode36%—
GSO—4.9%
AlgoTune1.41—

Agentic & Tool Use Kimi K2 (Jul 2025) leads

gpt-oss-120b: 12.2 (#153), Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-120bKimi K2 (Jul 2025)
Terminal-Bench18.7%35.7%
METR Time Horizons56.6%59.2%
APEX-Agents4.4%—
Berkeley Function Calling Leaderboard—59.1%
Vending-Bench 2-21.53—

Reasoning Kimi K2 (Jul 2025) leads

gpt-oss-120b: 20.0 (#245), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
Benchmarkgpt-oss-120bKimi K2 (Jul 2025)
SimpleBench22.1%26.3%
Kagi LLM Benchmark58.6%64.4%
LMArena Hard Prompts13641384
Epoch Capabilities Index139.93146.01
CritPt1.1%—
Chess Puzzles20%—
Mystery Game Puzzles2%—
DTBench76.3%—
LMCA22.1%—
Surface Evolver Bench25%—
ForecastBench—60.2

Math gpt-oss-120b leads

gpt-oss-120b: 52.5 (#50), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
Benchmarkgpt-oss-120bKimi K2 (Jul 2025)
Omni-MATH68.8%65.4%
LMArena Math13891397
OTIS Mock AIME 2024-202588.9%—
FrontierMath (Feb 2025 set)—21.4%
FrontierMath Tier 4 (v1)—0%

Knowledge gpt-oss-120b leads

gpt-oss-120b: 42.4 (#96), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
Benchmarkgpt-oss-120bKimi K2 (Jul 2025)
MMLU-Pro79.5%81.9%
Confabulations15.7%20.4%
Vectara Hallucination Rate14.2%17.9%
GPQA (HELM)68.4%65.3%
LMArena Expert13561365
GPQA Diamond75.8%—

Multilingual Kimi K2 (Jul 2025) leads

gpt-oss-120b: 48.0 (#147), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
Benchmarkgpt-oss-120bKimi K2 (Jul 2025)
LMArena Non-English13511372
LMArena Chinese13851415
LMArena French13691379
LMArena German13531387
LMArena Japanese13311349
LMArena Korean12821325
LMArena Russian13431385
LMArena Spanish13891386

Instruction Following Kimi K2 (Jul 2025) leads

gpt-oss-120b: 69.3 (#173), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
Benchmarkgpt-oss-120bKimi K2 (Jul 2025)
IFEval83.6%85%
LMArena Instruction Following13181348

Long Context Kimi K2 (Jul 2025) leads

gpt-oss-120b: 31.4 (#278), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
Benchmarkgpt-oss-120bKimi K2 (Jul 2025)
Fiction.LiveBench44.4%66.7%
LMArena Longer Query13191353
CL-bench—17.6%

Writing & Preference Kimi K2 (Jul 2025) leads

gpt-oss-120b: 46.5 (#217), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
Benchmarkgpt-oss-120bKimi K2 (Jul 2025)
LMArena Text13651380
LMArena Creative Writing12751350
Short-Story Creative Writing77.1%85.6%
EQ-Bench Creative Writing9611666
WildBench84.5%86.2%
LMArena Multi-Turn13401371

Frequently asked questions

Is gpt-oss-120b better than Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 36.3 on the Noometry Index. gpt-oss-120b costs 14× less per token, which makes it the better buy when Kimi K2 (Jul 2025)'s lead doesn't matter for your workload.

Which is cheaper, gpt-oss-120b or Kimi K2 (Jul 2025)?

gpt-oss-120b is cheaper. It lists at $0.037 per million input tokens and $0.17 per million output tokens; Kimi K2 (Jul 2025) lists at $0.57 and $2.30.

Is gpt-oss-120b or Kimi K2 (Jul 2025) better for coding?

Kimi K2 (Jul 2025) scores higher on coding benchmarks: 42.4 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

Kimi K2 (Jul 2025) does, with 262K tokens against 131K.

How many benchmarks do gpt-oss-120b and Kimi K2 (Jul 2025) share?

36 benchmarks have published results for both models. gpt-oss-120b has 48 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper