Model comparison

gpt-oss-120b vs Grok-2 (Dec 2024)

gpt-oss-120b is the stronger model overall, scoring 36.3 to 33.7 on the Noometry Index.

Last verified . 25 shared benchmarks.

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 25 benchmarks with published results for both. gpt-oss-120b scores higher in 6 categories and Grok-2 (Dec 2024) in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where gpt-oss-120b leads 52.5 to 20.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 88.9% for gpt-oss-120b and 11.5% for Grok-2 (Dec 2024).
  • gpt-oss-120b has downloadable open weights; the other is API-only.

Side by side

gpt-oss-120b and Grok-2 (Dec 2024) specifications
gpt-oss-120bGrok-2 (Dec 2024)
ProviderOpenAIxAI
Noometry Index36.333.7
Released2025-08-052024-08-13
WeightsOpenProprietary
Context window131K—
Max output41K—
Input $ / M tokens$0.037—
Output $ / M tokens$0.17—
Results tracked4834

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

gpt-oss-120b: 33.5 (#256), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
Benchmarkgpt-oss-120bGrok-2 (Dec 2024)
WeirdML48.2%22.2%
LMArena Coding13801287
SWE-bench Verified (bash only)26%—
Aider Polyglot41.8%—
SciCode36%—
LiveBench Coding—46.4%
ALE-Bench575.62—
AlgoTune1.41—

Agentic & Tool Use Not comparable

gpt-oss-120b: 12.2 (#153), Grok-2 (Dec 2024): —

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-120bGrok-2 (Dec 2024)
Terminal-Bench18.7%—
APEX-Agents4.4%—
METR Time Horizons56.6%—
Vending-Bench 2-21.53—

Reasoning gpt-oss-120b leads

gpt-oss-120b: 20.0 (#245), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
Benchmarkgpt-oss-120bGrok-2 (Dec 2024)
SimpleBench22.1%22.7%
LMArena Hard Prompts13641272
DTBench76.3%65.2%
Epoch Capabilities Index139.93130.48
Kagi LLM Benchmark58.6%—
CritPt1.1%—
Chess Puzzles20%—
LiveBench Reasoning—54.8%
Mystery Game Puzzles2%—
LiveBench Data Analysis—54.5%
LMCA22.1%—
Surface Evolver Bench25%—
LiveBench—54.3%

Math gpt-oss-120b leads

gpt-oss-120b: 52.5 (#50), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
Benchmarkgpt-oss-120bGrok-2 (Dec 2024)
OTIS Mock AIME 2024-202588.9%11.5%
LMArena Math13891283
Omni-MATH68.8%—
LiveBench Math—54.9%
MATH Level 5—63.5%
FrontierMath (Feb 2025 set)—0.7%

Knowledge gpt-oss-120b leads

gpt-oss-120b: 42.4 (#96), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
Benchmarkgpt-oss-120bGrok-2 (Dec 2024)
GPQA Diamond75.8%53.8%
Confabulations15.7%20.1%
LMArena Expert13561254
MMLU-Pro79.5%—
Vectara Hallucination Rate14.2%—
GPQA (HELM)68.4%—

Multilingual gpt-oss-120b leads

gpt-oss-120b: 48.0 (#147), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
Benchmarkgpt-oss-120bGrok-2 (Dec 2024)
LMArena Non-English13511282
LMArena Chinese13851289
LMArena French13691318
LMArena German13531287
LMArena Japanese13311244
LMArena Korean12821237
LMArena Russian13431286
LMArena Spanish13891281

Instruction Following gpt-oss-120b leads

gpt-oss-120b: 69.3 (#173), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
Benchmarkgpt-oss-120bGrok-2 (Dec 2024)
LMArena Instruction Following13181270
LiveBench Instruction Following—69.6%
IFEval83.6%—

Long Context Grok-2 (Dec 2024) leads

gpt-oss-120b: 31.4 (#278), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
Benchmarkgpt-oss-120bGrok-2 (Dec 2024)
LMArena Longer Query13191276
Fiction.LiveBench44.4%—

Writing & Preference Grok-2 (Dec 2024) leads

gpt-oss-120b: 46.5 (#217), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
Benchmarkgpt-oss-120bGrok-2 (Dec 2024)
LMArena Text13651305
LMArena Creative Writing12751284
Short-Story Creative Writing77.1%63.6%
LMArena Multi-Turn13401290
EQ-Bench Creative Writing961—
WildBench84.5%—
LiveBench Language—45.6%

Frequently asked questions

Is gpt-oss-120b better than Grok-2 (Dec 2024)?

gpt-oss-120b is the stronger model overall, scoring 36.3 to 33.7 on the Noometry Index.

Is gpt-oss-120b or Grok-2 (Dec 2024) better for coding?

They score almost the same on coding (33.5 vs 33.3); test both on your own repository before choosing.

How many benchmarks do gpt-oss-120b and Grok-2 (Dec 2024) share?

25 benchmarks have published results for both models. gpt-oss-120b has 48 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper