Model comparison

Grok 4 Fast vs Qwen3.6 Plus

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 39.4 on the Noometry Index.

Last verified . 20 shared benchmarks.

Grok 4 Fast xAI

39.4

Rank #167 Confirmed

Qwen3.6 Plus Alibaba (Qwen)

47.5

Rank #62 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Grok 4 Fast scores higher in 1 category and Qwen3.6 Plus in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.6 Plus leads 56.1 to 32.0.

Side by side

Grok 4 Fast and Qwen3.6 Plus specifications
Grok 4 FastQwen3.6 Plus
ProviderxAIAlibaba (Qwen)
Noometry Index39.447.5
Released2025-09-192026-03-31
WeightsProprietaryProprietary
Context window—1M
Max output—66K
Input $ / M tokens—$0.50
Output $ / M tokens—$3
Results tracked3037

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Plus leads

Grok 4 Fast: 32.5 (#271), Qwen3.6 Plus: 40.8 (#130)

Coding benchmarks
BenchmarkGrok 4 FastQwen3.6 Plus
LMArena WebDev11591461
LMArena Coding14291467
SWE-bench Verified—57.9%
SciCode—40.7%
WeirdML42.9%—
ALE-Bench—670.15

Agentic & Tool Use Not comparable

Grok 4 Fast: 29.5 (#86), Qwen3.6 Plus: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4 FastQwen3.6 Plus
τ²-bench Banking15.7%—
Cybench30%—
LMArena Search1171—
Vending-Bench 2—5,115

Reasoning Qwen3.6 Plus leads

Grok 4 Fast: 22.2 (#201), Qwen3.6 Plus: 29.3 (#93)

Reasoning benchmarks
BenchmarkGrok 4 FastQwen3.6 Plus
LMArena Hard Prompts14121449
DTBench82.7%81.9%
Epoch Capabilities Index144.2147.65
ARC-AGI-25.3%—
Kagi LLM Benchmark66.1%—
NYT Connections (extended)—60.3%
ARC-AGI-148.5%—
CritPt—2.9%
Chess Puzzles—17%
Thematic Generalization—59.5%
Mystery Game Puzzles—12%
LMCA—33.1%
ForecastBench60.5—

Math Qwen3.6 Plus leads

Grok 4 Fast: 38.9 (#123), Qwen3.6 Plus: 51.8 (#54)

Math benchmarks
BenchmarkGrok 4 FastQwen3.6 Plus
LMArena Math14191450
FrontierMath (Tiers 1-3)—38.2%
OTIS Mock AIME 2024-2025—93.3%
FrontierMath (Feb 2025 set)—26.2%
FrontierMath Tier 4 (v1)—8.3%

Knowledge Qwen3.6 Plus leads

Grok 4 Fast: 32.0 (#214), Qwen3.6 Plus: 56.1 (#45)

Knowledge benchmarks
BenchmarkGrok 4 FastQwen3.6 Plus
LMArena Expert14111454
GPQA Diamond—88.4%
SimpleQA Verified—44.1%
Vectara Hallucination Rate19.7%—

Multilingual Qwen3.6 Plus leads

Grok 4 Fast: 51.3 (#111), Qwen3.6 Plus: 53.3 (#70)

Multilingual benchmarks
BenchmarkGrok 4 FastQwen3.6 Plus
LMArena Non-English13961424
LMArena Chinese14571477
LMArena French14311455
LMArena German13831452
LMArena Japanese13521389
LMArena Korean13571379
LMArena Russian13891434
LMArena Spanish14151432

Instruction Following Qwen3.6 Plus leads

Grok 4 Fast: 73.2 (#121), Qwen3.6 Plus: 75.0 (#74)

Instruction Following benchmarks
BenchmarkGrok 4 FastQwen3.6 Plus
LMArena Instruction Following13871425

Long Context Grok 4 Fast leads

Grok 4 Fast: 63.2 (#3), Qwen3.6 Plus: 45.2 (#49)

Long Context benchmarks
BenchmarkGrok 4 FastQwen3.6 Plus
LMArena Longer Query14151439
Fiction.LiveBench94.4%—
CL-bench—20.3%

Writing & Preference Qwen3.6 Plus leads

Grok 4 Fast: 60.0 (#102), Qwen3.6 Plus: 62.2 (#82)

Writing & Preference benchmarks
BenchmarkGrok 4 FastQwen3.6 Plus
LMArena Text14071437
LMArena Creative Writing13871404
LMArena Multi-Turn14141438

Frequently asked questions

Is Grok 4 Fast better than Qwen3.6 Plus?

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 39.4 on the Noometry Index.

Is Grok 4 Fast or Qwen3.6 Plus better for coding?

Qwen3.6 Plus scores higher on coding benchmarks: 40.8 versus 32.5 in the Noometry coding category.

How many benchmarks do Grok 4 Fast and Qwen3.6 Plus share?

20 benchmarks have published results for both models. Grok 4 Fast has 30 scored results on Noometry and Qwen3.6 Plus has 37.

Related comparisons

Go deeper