Model comparison

Grok 4.1 Fast vs Qwen Max

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 34.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok 4.1 Fast scores higher in 8 categories and Qwen Max in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 25.1.
  • Grok 4.1 Fast is cheaper at $0.20 / $0.50 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Grok 4.1 Fast accepts more context: 128K tokens versus 33K.

Side by side

Grok 4.1 Fast and Qwen Max specifications
Grok 4.1 FastQwen Max
ProviderxAIAlibaba (Qwen)
Noometry Index41.434.7
Released2025-06-272024-04-03
WeightsProprietaryProprietary
Context window128K33K
Max output30K8K
Input $ / M tokens$0.20$1.60
Output $ / M tokens$0.50$6.40
Results tracked3223

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.1 Fast leads

Grok 4.1 Fast: 34.1 (#245), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkGrok 4.1 FastQwen Max
LMArena Coding14111288
Aider Polyglot—21.8%
LMArena WebDev1242—
ALE-Bench394.93—

Agentic & Tool Use Not comparable

Grok 4.1 Fast: 36.3 (#39), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1 FastQwen Max
Berkeley Function Calling Leaderboard69.6%—
τ²-bench Banking13.1%—
LMArena Search1171—
Vending-Bench 21,107—

Reasoning Grok 4.1 Fast leads

Grok 4.1 Fast: 43.4 (#49), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkGrok 4.1 FastQwen Max
LMArena Hard Prompts14071269
SimpleBench56%—
NYT Connections (extended)87.4%—
DTBench87.7%—
ForecastBench61—

Math Grok 4.1 Fast leads

Grok 4.1 Fast: 31.9 (#221), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkGrok 4.1 FastQwen Max
LMArena Math14081275
MathArena Final-Answer Competitions60.9%—
OTIS Mock AIME 2024-2025—16.1%
ProofBench4%—
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Grok 4.1 Fast leads

Grok 4.1 Fast: 33.1 (#207), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkGrok 4.1 FastQwen Max
LMArena Expert13991248
GPQA Diamond—56.1%
Vectara Hallucination Rate17.8%—

Multimodal Not comparable

Grok 4.1 Fast: 37.0 (#76), Qwen Max: —

Multimodal benchmarks
BenchmarkGrok 4.1 FastQwen Max
LMArena Vision1201—

Multilingual Grok 4.1 Fast leads

Grok 4.1 Fast: 51.0 (#114), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkGrok 4.1 FastQwen Max
LMArena Non-English13911263
LMArena Chinese14411254
LMArena French14151330
LMArena German14041254
LMArena Japanese13491205
LMArena Korean13611142
LMArena Russian13871274
LMArena Spanish14131290

Instruction Following Grok 4.1 Fast leads

Grok 4.1 Fast: 72.7 (#133), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkGrok 4.1 FastQwen Max
LMArena Instruction Following13761262

Long Context Grok 4.1 Fast leads

Grok 4.1 Fast: 42.4 (#126), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkGrok 4.1 FastQwen Max
LMArena Longer Query13901288
Fiction.LiveBench—66.7%

Writing & Preference Grok 4.1 Fast leads

Grok 4.1 Fast: 57.2 (#131), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkGrok 4.1 FastQwen Max
LMArena Text14081282
LMArena Creative Writing13941248
LMArena Multi-Turn13891277
EQ-Bench Creative Writing1327—

Frequently asked questions

Is Grok 4.1 Fast better than Qwen Max?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 34.7 on the Noometry Index.

Which is cheaper, Grok 4.1 Fast or Qwen Max?

Grok 4.1 Fast is cheaper. It lists at $0.20 per million input tokens and $0.50 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Grok 4.1 Fast or Qwen Max better for coding?

Grok 4.1 Fast scores higher on coding benchmarks: 34.1 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Grok 4.1 Fast does, with 128K tokens against 33K.

How many benchmarks do Grok 4.1 Fast and Qwen Max share?

17 benchmarks have published results for both models. Grok 4.1 Fast has 32 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper