Model comparison

Grok 4 Fast vs Qwen3.5 Plus

Qwen3.5 Plus is the stronger model overall, scoring 42.9 to 39.4 on the Noometry Index.

Last verified . 3 shared benchmarks.

Grok 4 Fast xAI

39.4

Rank #167 Confirmed

Qwen3.5 Plus Alibaba (Qwen)

42.9

Rank #106 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Grok 4 Fast scores higher in 1 category and Qwen3.5 Plus in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Grok 4 Fast leads 63.2 to 43.0.
  • The biggest single-benchmark swing is Vectara Hallucination Rate: 19.7% for Grok 4 Fast and 10.7% for Qwen3.5 Plus.

Side by side

Grok 4 Fast and Qwen3.5 Plus specifications
Grok 4 FastQwen3.5 Plus
ProviderxAIAlibaba (Qwen)
Noometry Index39.442.9
Released2025-09-192026-02-16
WeightsProprietaryProprietary
Context window—1M
Max output—66K
Input $ / M tokens—$0.40
Output $ / M tokens—$2.40
Results tracked3015

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Grok 4 Fast: 32.5 (#271), Qwen3.5 Plus: —

Coding benchmarks
BenchmarkGrok 4 FastQwen3.5 Plus
LMArena WebDev1159—
WeirdML42.9%—
LMArena Coding1429—
ALE-Bench—621.92

Agentic & Tool Use Not comparable

Grok 4 Fast: 29.5 (#86), Qwen3.5 Plus: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4 FastQwen3.5 Plus
τ²-bench Banking15.7%—
Cybench30%—
LMArena Search1171—
Vending-Bench 2—0.54

Reasoning Qwen3.5 Plus leads

Grok 4 Fast: 22.2 (#201), Qwen3.5 Plus: 32.8 (#74)

Reasoning benchmarks
BenchmarkGrok 4 FastQwen3.5 Plus
DTBench82.7%80.5%
Epoch Capabilities Index144.2146.78
ARC-AGI-25.3%—
Kagi LLM Benchmark66.1%—
ARC-AGI-148.5%—
Chess Puzzles—22%
LMArena Hard Prompts1412—
Mystery Game Puzzles—17%
LMCA—36.4%
ForecastBench60.5—

Math Qwen3.5 Plus leads

Grok 4 Fast: 38.9 (#123), Qwen3.5 Plus: 49.6 (#61)

Math benchmarks
BenchmarkGrok 4 FastQwen3.5 Plus
OTIS Mock AIME 2024-2025—86.7%
LMArena Math1419—
FrontierMath (Feb 2025 set)—21%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Qwen3.5 Plus leads

Grok 4 Fast: 32.0 (#214), Qwen3.5 Plus: 46.0 (#83)

Knowledge benchmarks
BenchmarkGrok 4 FastQwen3.5 Plus
Vectara Hallucination Rate19.7%10.7%
GPQA Diamond—84.8%
SimpleQA Verified—25.4%
LMArena Expert1411—

Multilingual Not comparable

Grok 4 Fast: 51.3 (#111), Qwen3.5 Plus: —

Multilingual benchmarks
BenchmarkGrok 4 FastQwen3.5 Plus
LMArena Non-English1396—
LMArena Chinese1457—
LMArena French1431—
LMArena German1383—
LMArena Japanese1352—
LMArena Korean1357—
LMArena Russian1389—
LMArena Spanish1415—

Instruction Following Not comparable

Grok 4 Fast: 73.2 (#121), Qwen3.5 Plus: —

Instruction Following benchmarks
BenchmarkGrok 4 FastQwen3.5 Plus
LMArena Instruction Following1387—

Long Context Grok 4 Fast leads

Grok 4 Fast: 63.2 (#3), Qwen3.5 Plus: 43.0 (#113)

Long Context benchmarks
BenchmarkGrok 4 FastQwen3.5 Plus
Fiction.LiveBench94.4%—
CL-bench—19.8%
CL-bench Life—12.4%
LMArena Longer Query1415—

Writing & Preference Not comparable

Grok 4 Fast: 60.0 (#102), Qwen3.5 Plus: —

Writing & Preference benchmarks
BenchmarkGrok 4 FastQwen3.5 Plus
LMArena Text1407—
LMArena Creative Writing1387—
LMArena Multi-Turn1414—

Frequently asked questions

Is Grok 4 Fast better than Qwen3.5 Plus?

Qwen3.5 Plus is the stronger model overall, scoring 42.9 to 39.4 on the Noometry Index.

How many benchmarks do Grok 4 Fast and Qwen3.5 Plus share?

3 benchmarks have published results for both models. Grok 4 Fast has 30 scored results on Noometry and Qwen3.5 Plus has 15.

Related comparisons

Go deeper