Model comparison

Grok 4.1 Fast vs Qwen3.5-9B

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 33.8 on the Noometry Index. Qwen3.5-9B costs 2.4× less per token, which makes it the better buy when Grok 4.1 Fast's lead doesn't matter for your workload.

Last verified . 2 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Qwen3.5-9B Alibaba (Qwen)

33.8

Rank #236 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Grok 4.1 Fast scores higher in 2 categories and Qwen3.5-9B in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Grok 4.1 Fast leads 36.3 to 14.5.
  • The biggest single-benchmark swing is DTBench: 87.7% for Grok 4.1 Fast and 71.2% for Qwen3.5-9B.
  • Qwen3.5-9B is cheaper at $0.10 / $0.15 per million input/output tokens, against $0.20 / $0.50 for Grok 4.1 Fast.
  • Qwen3.5-9B accepts more context: 262K tokens versus 128K.
  • Qwen3.5-9B has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 Fast and Qwen3.5-9B specifications
Grok 4.1 FastQwen3.5-9B
ProviderxAIAlibaba (Qwen)
Noometry Index41.433.8
Released2025-06-272026-02-23
WeightsProprietaryOpen
Context window128K262K
Max output30K66K
Input $ / M tokens$0.20$0.10
Output $ / M tokens$0.50$0.15
Results tracked3210

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5-9B leads

Grok 4.1 Fast: 34.1 (#245), Qwen3.5-9B: 35.9 (#217)

Coding benchmarks
BenchmarkGrok 4.1 FastQwen3.5-9B
LMArena WebDev1242—
SciCode—27.5%
LMArena Coding1411—
ALE-Bench394.93—

Agentic & Tool Use Grok 4.1 Fast leads

Grok 4.1 Fast: 36.3 (#39), Qwen3.5-9B: 14.5 (#151)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1 FastQwen3.5-9B
Terminal-Bench—9.2%
Berkeley Function Calling Leaderboard69.6%—
τ²-bench Banking13.1%—
LMArena Search1171—
Vending-Bench 21,107—

Reasoning Grok 4.1 Fast leads

Grok 4.1 Fast: 43.4 (#49), Qwen3.5-9B: 23.1 (#182)

Reasoning benchmarks
BenchmarkGrok 4.1 FastQwen3.5-9B
DTBench87.7%71.2%
SimpleBench56%—
NYT Connections (extended)87.4%—
CritPt—0.3%
Chess Puzzles—12%
LMArena Hard Prompts1407—
LMCA—24.5%
Epoch Capabilities Index—139.46
ForecastBench61—

Math Qwen3.5-9B leads

Grok 4.1 Fast: 31.9 (#221), Qwen3.5-9B: 34.8 (#192)

Math benchmarks
BenchmarkGrok 4.1 FastQwen3.5-9B
MathArena Final-Answer Competitions60.9%48.5%
OTIS Mock AIME 2024-2025—61.7%
ProofBench4%—
LMArena Math1408—

Knowledge Qwen3.5-9B leads

Grok 4.1 Fast: 33.1 (#207), Qwen3.5-9B: 46.0 (#84)

Knowledge benchmarks
BenchmarkGrok 4.1 FastQwen3.5-9B
GPQA Diamond—79%
Vectara Hallucination Rate17.8%—
LMArena Expert1399—

Multimodal Not comparable

Grok 4.1 Fast: 37.0 (#76), Qwen3.5-9B: —

Multimodal benchmarks
BenchmarkGrok 4.1 FastQwen3.5-9B
LMArena Vision1201—

Multilingual Not comparable

Grok 4.1 Fast: 51.0 (#114), Qwen3.5-9B: —

Multilingual benchmarks
BenchmarkGrok 4.1 FastQwen3.5-9B
LMArena Non-English1391—
LMArena Chinese1441—
LMArena French1415—
LMArena German1404—
LMArena Japanese1349—
LMArena Korean1361—
LMArena Russian1387—
LMArena Spanish1413—

Instruction Following Not comparable

Grok 4.1 Fast: 72.7 (#133), Qwen3.5-9B: —

Instruction Following benchmarks
BenchmarkGrok 4.1 FastQwen3.5-9B
LMArena Instruction Following1376—

Long Context Not comparable

Grok 4.1 Fast: 42.4 (#126), Qwen3.5-9B: —

Long Context benchmarks
BenchmarkGrok 4.1 FastQwen3.5-9B
LMArena Longer Query1390—

Writing & Preference Not comparable

Grok 4.1 Fast: 57.2 (#131), Qwen3.5-9B: —

Writing & Preference benchmarks
BenchmarkGrok 4.1 FastQwen3.5-9B
LMArena Text1408—
LMArena Creative Writing1394—
EQ-Bench Creative Writing1327—
LMArena Multi-Turn1389—

Frequently asked questions

Is Grok 4.1 Fast better than Qwen3.5-9B?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 33.8 on the Noometry Index. Qwen3.5-9B costs 2.4× less per token, which makes it the better buy when Grok 4.1 Fast's lead doesn't matter for your workload.

Which is cheaper, Grok 4.1 Fast or Qwen3.5-9B?

Qwen3.5-9B is cheaper. It lists at $0.10 per million input tokens and $0.15 per million output tokens; Grok 4.1 Fast lists at $0.20 and $0.50.

Is Grok 4.1 Fast or Qwen3.5-9B better for coding?

Qwen3.5-9B scores higher on coding benchmarks: 35.9 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Qwen3.5-9B does, with 262K tokens against 128K.

How many benchmarks do Grok 4.1 Fast and Qwen3.5-9B share?

2 benchmarks have published results for both models. Grok 4.1 Fast has 32 scored results on Noometry and Qwen3.5-9B has 10.

Related comparisons

Go deeper