Model comparison

Grok Build 0.1 vs Qwen2.5-Coder-32B

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 33.4 on the Noometry Index. Qwen2.5-Coder-32B costs 1.7× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Summary

  • The widest gap is in coding, where Grok Build 0.1 leads 43.1 to 22.6.
  • Qwen2.5-Coder-32B is cheaper at $0.66 / $1 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Grok Build 0.1 accepts more context: 256K tokens versus 33K.
  • Qwen2.5-Coder-32B has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Qwen2.5-Coder-32B specifications
Grok Build 0.1Qwen2.5-Coder-32B
ProviderxAIAlibaba (Qwen)
Noometry Index36.433.4
Released2026-04-162024-09-18
WeightsProprietaryOpen
Context window256K33K
Max output256K29K
Input $ / M tokens$1$0.66
Output $ / M tokens$2$1
Results tracked331

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Qwen2.5-Coder-32B: 22.6 (#333)

Coding benchmarks
BenchmarkGrok Build 0.1Qwen2.5-Coder-32B
SWE-bench Verified (bash only)—9%
Aider Polyglot—16.4%
SciCode50.2%—
BigCodeBench Instruct—49%
LiveBench Coding—56.9%
LMArena Coding—1276
BigCodeBench Complete—58%
HumanEval+—87.2%
MBPP+—77%

Agentic & Tool Use Not comparable

Grok Build 0.1: 22.7 (#129), Qwen2.5-Coder-32B: —

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Qwen2.5-Coder-32B
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Qwen2.5-Coder-32B: 21.2 (#225)

Reasoning benchmarks
BenchmarkGrok Build 0.1Qwen2.5-Coder-32B
CritPt9.1%—
LiveBench Reasoning—42.1%
LMArena Hard Prompts—1251
LiveBench Data Analysis—49.9%
Epoch Capabilities Index—119.49
HellaSwag—83%
LiveBench—46.2%
WinoGrande—80.8%

Math Not comparable

Grok Build 0.1: —, Qwen2.5-Coder-32B: 33.3 (#204)

Math benchmarks
BenchmarkGrok Build 0.1Qwen2.5-Coder-32B
LiveBench Math—46.6%
LMArena Math—1251
GSM8K—93%

Knowledge Not comparable

Grok Build 0.1: —, Qwen2.5-Coder-32B: 33.4 (#203)

Knowledge benchmarks
BenchmarkGrok Build 0.1Qwen2.5-Coder-32B
LMArena Expert—1221
ARC (AI2) Challenge—70.5%
MMLU—79.1%

Multilingual Not comparable

Grok Build 0.1: —, Qwen2.5-Coder-32B: 37.8 (#235)

Multilingual benchmarks
BenchmarkGrok Build 0.1Qwen2.5-Coder-32B
LMArena Non-English—1205
LMArena Chinese—1222
LMArena Russian—1228

Instruction Following Not comparable

Grok Build 0.1: —, Qwen2.5-Coder-32B: 61.4 (#245)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Qwen2.5-Coder-32B
LiveBench Instruction Following—58.7%
LMArena Instruction Following—1223

Long Context Not comparable

Grok Build 0.1: —, Qwen2.5-Coder-32B: 38.0 (#208)

Long Context benchmarks
BenchmarkGrok Build 0.1Qwen2.5-Coder-32B
LMArena Longer Query—1251

Writing & Preference Not comparable

Grok Build 0.1: —, Qwen2.5-Coder-32B: 41.6 (#240)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Qwen2.5-Coder-32B
LMArena Text—1230
LMArena Creative Writing—1174
LMArena Multi-Turn—1222
LiveBench Language—23.3%

Frequently asked questions

Is Grok Build 0.1 better than Qwen2.5-Coder-32B?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 33.4 on the Noometry Index. Qwen2.5-Coder-32B costs 1.7× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.

Which is cheaper, Grok Build 0.1 or Qwen2.5-Coder-32B?

Qwen2.5-Coder-32B is cheaper. It lists at $0.66 per million input tokens and $1 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or Qwen2.5-Coder-32B better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 22.6 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 33K.

How many benchmarks do Grok Build 0.1 and Qwen2.5-Coder-32B share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Qwen2.5-Coder-32B has 31.

Related comparisons

Go deeper