Model comparison

Grok 4 Fast vs Qwen2.5 Plus 1127

Grok 4 Fast and Qwen2.5 Plus 1127 score almost the same on the Noometry Index (39.4 vs 38.8), so choose on price, context window or the category you care about most.

Last verified . 14 shared benchmarks.

Grok 4 Fast xAI

39.4

Rank #167 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Grok 4 Fast scores higher in 5 categories and Qwen2.5 Plus 1127 in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Grok 4 Fast leads 63.2 to 39.2.

Side by side

Grok 4 Fast and Qwen2.5 Plus 1127 specifications
Grok 4 FastQwen2.5 Plus 1127
ProviderxAIAlibaba (Qwen)
Noometry Index39.438.8
Released2025-09-19—
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3014

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Grok 4 Fast: 32.5 (#271), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkGrok 4 FastQwen2.5 Plus 1127
LMArena Coding14291314
LMArena WebDev1159—
WeirdML42.9%—

Agentic & Tool Use Not comparable

Grok 4 Fast: 29.5 (#86), Qwen2.5 Plus 1127: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4 FastQwen2.5 Plus 1127
τ²-bench Banking15.7%—
Cybench30%—
LMArena Search1171—

Reasoning Qwen2.5 Plus 1127 leads

Grok 4 Fast: 22.2 (#201), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkGrok 4 FastQwen2.5 Plus 1127
LMArena Hard Prompts14121299
ARC-AGI-25.3%—
Kagi LLM Benchmark66.1%—
ARC-AGI-148.5%—
DTBench82.7%—
Epoch Capabilities Index144.2—
ForecastBench60.5—

Math Grok 4 Fast leads

Grok 4 Fast: 38.9 (#123), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkGrok 4 FastQwen2.5 Plus 1127
LMArena Math14191298

Knowledge Qwen2.5 Plus 1127 leads

Grok 4 Fast: 32.0 (#214), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkGrok 4 FastQwen2.5 Plus 1127
LMArena Expert14111289
Vectara Hallucination Rate19.7%—

Multilingual Grok 4 Fast leads

Grok 4 Fast: 51.3 (#111), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkGrok 4 FastQwen2.5 Plus 1127
LMArena Non-English13961265
LMArena Chinese14571314
LMArena German13831231
LMArena Japanese13521207
LMArena Russian13891271
LMArena French1431—
LMArena Korean1357—
LMArena Spanish1415—

Instruction Following Grok 4 Fast leads

Grok 4 Fast: 73.2 (#121), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkGrok 4 FastQwen2.5 Plus 1127
LMArena Instruction Following13871275

Long Context Grok 4 Fast leads

Grok 4 Fast: 63.2 (#3), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkGrok 4 FastQwen2.5 Plus 1127
LMArena Longer Query14151292
Fiction.LiveBench94.4%—

Writing & Preference Grok 4 Fast leads

Grok 4 Fast: 60.0 (#102), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkGrok 4 FastQwen2.5 Plus 1127
LMArena Text14071299
LMArena Creative Writing13871262
LMArena Multi-Turn14141299

Frequently asked questions

Is Grok 4 Fast better than Qwen2.5 Plus 1127?

Grok 4 Fast and Qwen2.5 Plus 1127 score almost the same on the Noometry Index (39.4 vs 38.8), so choose on price, context window or the category you care about most.

Is Grok 4 Fast or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 32.5 in the Noometry coding category.

How many benchmarks do Grok 4 Fast and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. Grok 4 Fast has 30 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper