Model comparison

Grok 4.1 Fast vs Qwen3 Max

Qwen3 Max is the stronger model overall, scoring 43.7 to 41.4 on the Noometry Index. Grok 4.1 Fast costs 8.7× less per token, which makes it the better buy when Qwen3 Max's lead doesn't matter for your workload.

Last verified . 21 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Grok 4.1 Fast scores higher in 2 categories and Qwen3 Max in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 22.6.
  • The biggest single-benchmark swing is NYT Connections (extended): 87.4% for Grok 4.1 Fast and 30.1% for Qwen3 Max.
  • Grok 4.1 Fast is cheaper at $0.20 / $0.50 per million input/output tokens, against $1.20 / $6 for Qwen3 Max.
  • Qwen3 Max accepts more context: 262K tokens versus 128K.

Side by side

Grok 4.1 Fast and Qwen3 Max specifications
Grok 4.1 FastQwen3 Max
ProviderxAIAlibaba (Qwen)
Noometry Index41.443.7
Released2025-06-272025-09-23
WeightsProprietaryProprietary
Context window128K262K
Max output30K66K
Input $ / M tokens$0.20$1.20
Output $ / M tokens$0.50$6
Results tracked3233

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 Max leads

Grok 4.1 Fast: 34.1 (#245), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkGrok 4.1 FastQwen3 Max
LMArena Coding14111456
ALE-Bench394.93370.45
LMArena WebDev1242—

Agentic & Tool Use Not comparable

Grok 4.1 Fast: 36.3 (#39), Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1 FastQwen3 Max
Vending-Bench 21,10771.56
Berkeley Function Calling Leaderboard69.6%—
τ²-bench Banking13.1%—
LMArena Search1171—

Reasoning Grok 4.1 Fast leads

Grok 4.1 Fast: 43.4 (#49), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkGrok 4.1 FastQwen3 Max
NYT Connections (extended)87.4%30.1%
LMArena Hard Prompts14071448
DTBench87.7%82.1%
SimpleBench56%—
Kagi LLM Benchmark—72.5%
Chess Puzzles—4%
Mystery Game Puzzles—5%
LMCA—28.3%
Epoch Capabilities Index—142.38
ForecastBench61—

Math Qwen3 Max leads

Grok 4.1 Fast: 31.9 (#221), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkGrok 4.1 FastQwen3 Max
LMArena Math14081446
FrontierMath (Tiers 1-3)—18.9%
MathArena Final-Answer Competitions60.9%—
OTIS Mock AIME 2024-2025—73.3%
ProofBench4%—
MATH Level 5—97.1%

Knowledge Qwen3 Max leads

Grok 4.1 Fast: 33.1 (#207), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkGrok 4.1 FastQwen3 Max
LMArena Expert13991455
GPQA Diamond—72.6%
SimpleQA Verified—48.7%
Vectara Hallucination Rate17.8%—

Multimodal Not comparable

Grok 4.1 Fast: 37.0 (#76), Qwen3 Max: —

Multimodal benchmarks
BenchmarkGrok 4.1 FastQwen3 Max
LMArena Vision1201—

Multilingual Qwen3 Max leads

Grok 4.1 Fast: 51.0 (#114), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkGrok 4.1 FastQwen3 Max
LMArena Non-English13911429
LMArena Chinese14411478
LMArena French14151449
LMArena German14041463
LMArena Japanese13491397
LMArena Korean13611399
LMArena Russian13871428
LMArena Spanish14131462

Instruction Following Qwen3 Max leads

Grok 4.1 Fast: 72.7 (#133), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkGrok 4.1 FastQwen3 Max
LMArena Instruction Following13761419

Long Context Too close to call

Grok 4.1 Fast: 42.4 (#126), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkGrok 4.1 FastQwen3 Max
LMArena Longer Query13901438
Fiction.LiveBench—66.7%
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

Grok 4.1 Fast: 57.2 (#131), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkGrok 4.1 FastQwen3 Max
LMArena Text14081439
LMArena Creative Writing13941402
LMArena Multi-Turn13891446
EQ-Bench Creative Writing1327—

Frequently asked questions

Is Grok 4.1 Fast better than Qwen3 Max?

Qwen3 Max is the stronger model overall, scoring 43.7 to 41.4 on the Noometry Index. Grok 4.1 Fast costs 8.7× less per token, which makes it the better buy when Qwen3 Max's lead doesn't matter for your workload.

Which is cheaper, Grok 4.1 Fast or Qwen3 Max?

Grok 4.1 Fast is cheaper. It lists at $0.20 per million input tokens and $0.50 per million output tokens; Qwen3 Max lists at $1.20 and $6.

Is Grok 4.1 Fast or Qwen3 Max better for coding?

Qwen3 Max scores higher on coding benchmarks: 43.0 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Qwen3 Max does, with 262K tokens against 128K.

How many benchmarks do Grok 4.1 Fast and Qwen3 Max share?

21 benchmarks have published results for both models. Grok 4.1 Fast has 32 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper