Model comparison

Grok-2 (Dec 2024) vs Qwen2.5-Coder-32B

Grok-2 (Dec 2024) and Qwen2.5-Coder-32B score almost the same on the Noometry Index (33.7 vs 33.4), so choose on price, context window or the category you care about most.

Last verified . 20 shared benchmarks.

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Grok-2 (Dec 2024) scores higher in 5 categories and Qwen2.5-Coder-32B in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen2.5-Coder-32B leads 33.3 to 20.8.
  • The biggest single-benchmark swing is LiveBench Language: 45.6% for Grok-2 (Dec 2024) and 23.3% for Qwen2.5-Coder-32B.
  • Qwen2.5-Coder-32B has downloadable open weights; the other is API-only.

Side by side

Grok-2 (Dec 2024) and Qwen2.5-Coder-32B specifications
Grok-2 (Dec 2024)Qwen2.5-Coder-32B
ProviderxAIAlibaba (Qwen)
Noometry Index33.733.4
Released2024-08-132024-09-18
WeightsProprietaryOpen
Context window—33K
Max output—29K
Input $ / M tokens—$0.66
Output $ / M tokens—$1
Results tracked3431

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok-2 (Dec 2024) leads

Grok-2 (Dec 2024): 33.3 (#258), Qwen2.5-Coder-32B: 22.6 (#333)

Coding benchmarks
BenchmarkGrok-2 (Dec 2024)Qwen2.5-Coder-32B
LiveBench Coding46.4%56.9%
LMArena Coding12871276
SWE-bench Verified (bash only)—9%
Aider Polyglot—16.4%
WeirdML22.2%—
BigCodeBench Instruct—49%
BigCodeBench Complete—58%
HumanEval+—87.2%
MBPP+—77%

Reasoning Qwen2.5-Coder-32B leads

Grok-2 (Dec 2024): 16.9 (#299), Qwen2.5-Coder-32B: 21.2 (#225)

Reasoning benchmarks
BenchmarkGrok-2 (Dec 2024)Qwen2.5-Coder-32B
LiveBench Reasoning54.8%42.1%
LMArena Hard Prompts12721251
LiveBench Data Analysis54.5%49.9%
Epoch Capabilities Index130.48119.49
LiveBench54.3%46.2%
SimpleBench22.7%—
DTBench65.2%—
HellaSwag—83%
WinoGrande—80.8%

Math Qwen2.5-Coder-32B leads

Grok-2 (Dec 2024): 20.8 (#284), Qwen2.5-Coder-32B: 33.3 (#204)

Math benchmarks
BenchmarkGrok-2 (Dec 2024)Qwen2.5-Coder-32B
LiveBench Math54.9%46.6%
LMArena Math12831251
OTIS Mock AIME 2024-202511.5%—
MATH Level 563.5%—
FrontierMath (Feb 2025 set)0.7%—
GSM8K—93%

Knowledge Qwen2.5-Coder-32B leads

Grok-2 (Dec 2024): 29.8 (#233), Qwen2.5-Coder-32B: 33.4 (#203)

Knowledge benchmarks
BenchmarkGrok-2 (Dec 2024)Qwen2.5-Coder-32B
LMArena Expert12541221
GPQA Diamond53.8%—
Confabulations20.1%—
ARC (AI2) Challenge—70.5%
MMLU—79.1%

Multilingual Grok-2 (Dec 2024) leads

Grok-2 (Dec 2024): 43.1 (#188), Qwen2.5-Coder-32B: 37.8 (#235)

Multilingual benchmarks
BenchmarkGrok-2 (Dec 2024)Qwen2.5-Coder-32B
LMArena Non-English12821205
LMArena Chinese12891222
LMArena Russian12861228
LMArena French1318—
LMArena German1287—
LMArena Japanese1244—
LMArena Korean1237—
LMArena Spanish1281—

Instruction Following Grok-2 (Dec 2024) leads

Grok-2 (Dec 2024): 66.9 (#202), Qwen2.5-Coder-32B: 61.4 (#245)

Instruction Following benchmarks
BenchmarkGrok-2 (Dec 2024)Qwen2.5-Coder-32B
LiveBench Instruction Following69.6%58.7%
LMArena Instruction Following12701223

Long Context Too close to call

Grok-2 (Dec 2024): 38.8 (#190), Qwen2.5-Coder-32B: 38.0 (#208)

Long Context benchmarks
BenchmarkGrok-2 (Dec 2024)Qwen2.5-Coder-32B
LMArena Longer Query12761251

Writing & Preference Grok-2 (Dec 2024) leads

Grok-2 (Dec 2024): 48.6 (#198), Qwen2.5-Coder-32B: 41.6 (#240)

Writing & Preference benchmarks
BenchmarkGrok-2 (Dec 2024)Qwen2.5-Coder-32B
LMArena Text13051230
LMArena Creative Writing12841174
LMArena Multi-Turn12901222
LiveBench Language45.6%23.3%
Short-Story Creative Writing63.6%—

Frequently asked questions

Is Grok-2 (Dec 2024) better than Qwen2.5-Coder-32B?

Grok-2 (Dec 2024) and Qwen2.5-Coder-32B score almost the same on the Noometry Index (33.7 vs 33.4), so choose on price, context window or the category you care about most.

Is Grok-2 (Dec 2024) or Qwen2.5-Coder-32B better for coding?

Grok-2 (Dec 2024) scores higher on coding benchmarks: 33.3 versus 22.6 in the Noometry coding category.

How many benchmarks do Grok-2 (Dec 2024) and Qwen2.5-Coder-32B share?

20 benchmarks have published results for both models. Grok-2 (Dec 2024) has 34 scored results on Noometry and Qwen2.5-Coder-32B has 31.

Related comparisons

Go deeper