Model comparison

Grok 4.1 Fast vs Qwen3-Coder 480B-A35B Instruct

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 38.1 on the Noometry Index.

Last verified . 19 shared benchmarks.

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Grok 4.1 Fast scores higher in 6 categories and Qwen3-Coder 480B-A35B Instruct in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 25.5.
  • Grok 4.1 Fast is cheaper at $0.20 / $0.50 per million input/output tokens, against $1.50 / $7.50 for Qwen3-Coder 480B-A35B Instruct.
  • Qwen3-Coder 480B-A35B Instruct accepts more context: 262K tokens versus 128K.
  • Qwen3-Coder 480B-A35B Instruct has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 Fast and Qwen3-Coder 480B-A35B Instruct specifications
Grok 4.1 FastQwen3-Coder 480B-A35B Instruct
ProviderxAIAlibaba (Qwen)
Noometry Index41.438.1
Released2025-06-272025-04
WeightsProprietaryOpen
Context window128K262K
Max output30K66K
Input $ / M tokens$0.20$1.50
Output $ / M tokens$0.50$7.50
Results tracked3225

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3-Coder 480B-A35B Instruct leads

Grok 4.1 Fast: 34.1 (#245), Qwen3-Coder 480B-A35B Instruct: 35.5 (#223)

Coding benchmarks
BenchmarkGrok 4.1 FastQwen3-Coder 480B-A35B Instruct
LMArena WebDev12421275
LMArena Coding14111412
ALE-Bench394.93461.45
SWE-bench Verified (bash only)—55.4%
GSO—4.9%
WeirdML—41.2%
AlgoTune—1.44

Agentic & Tool Use Grok 4.1 Fast leads

Grok 4.1 Fast: 36.3 (#39), Qwen3-Coder 480B-A35B Instruct: 23.9 (#123)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1 FastQwen3-Coder 480B-A35B Instruct
Terminal-Bench—27.2%
Berkeley Function Calling Leaderboard69.6%—
τ²-bench Banking13.1%—
LMArena Search1171—
Vending-Bench 21,107—

Reasoning Grok 4.1 Fast leads

Grok 4.1 Fast: 43.4 (#49), Qwen3-Coder 480B-A35B Instruct: 25.5 (#149)

Reasoning benchmarks
BenchmarkGrok 4.1 FastQwen3-Coder 480B-A35B Instruct
LMArena Hard Prompts14071372
SimpleBench56%—
Kagi LLM Benchmark—49.5%
NYT Connections (extended)87.4%—
DTBench87.7%—
ForecastBench61—

Math Qwen3-Coder 480B-A35B Instruct leads

Grok 4.1 Fast: 31.9 (#221), Qwen3-Coder 480B-A35B Instruct: 37.6 (#150)

Math benchmarks
BenchmarkGrok 4.1 FastQwen3-Coder 480B-A35B Instruct
LMArena Math14081365
MathArena Final-Answer Competitions60.9%—
ProofBench4%—

Knowledge Qwen3-Coder 480B-A35B Instruct leads

Grok 4.1 Fast: 33.1 (#207), Qwen3-Coder 480B-A35B Instruct: 37.0 (#162)

Knowledge benchmarks
BenchmarkGrok 4.1 FastQwen3-Coder 480B-A35B Instruct
LMArena Expert13991338
Vectara Hallucination Rate17.8%—

Multimodal Not comparable

Grok 4.1 Fast: 37.0 (#76), Qwen3-Coder 480B-A35B Instruct: —

Multimodal benchmarks
BenchmarkGrok 4.1 FastQwen3-Coder 480B-A35B Instruct
LMArena Vision1201—

Multilingual Grok 4.1 Fast leads

Grok 4.1 Fast: 51.0 (#114), Qwen3-Coder 480B-A35B Instruct: 47.7 (#148)

Multilingual benchmarks
BenchmarkGrok 4.1 FastQwen3-Coder 480B-A35B Instruct
LMArena Non-English13911346
LMArena Chinese14411357
LMArena French14151398
LMArena German14041325
LMArena Japanese13491310
LMArena Korean13611305
LMArena Russian13871366
LMArena Spanish14131360

Instruction Following Grok 4.1 Fast leads

Grok 4.1 Fast: 72.7 (#133), Qwen3-Coder 480B-A35B Instruct: 71.6 (#147)

Instruction Following benchmarks
BenchmarkGrok 4.1 FastQwen3-Coder 480B-A35B Instruct
LMArena Instruction Following13761355

Long Context Too close to call

Grok 4.1 Fast: 42.4 (#126), Qwen3-Coder 480B-A35B Instruct: 42.0 (#131)

Long Context benchmarks
BenchmarkGrok 4.1 FastQwen3-Coder 480B-A35B Instruct
LMArena Longer Query13901378

Writing & Preference Grok 4.1 Fast leads

Grok 4.1 Fast: 57.2 (#131), Qwen3-Coder 480B-A35B Instruct: 55.3 (#147)

Writing & Preference benchmarks
BenchmarkGrok 4.1 FastQwen3-Coder 480B-A35B Instruct
LMArena Text14081357
LMArena Creative Writing13941333
LMArena Multi-Turn13891365
EQ-Bench Creative Writing1327—

Frequently asked questions

Is Grok 4.1 Fast better than Qwen3-Coder 480B-A35B Instruct?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 38.1 on the Noometry Index.

Which is cheaper, Grok 4.1 Fast or Qwen3-Coder 480B-A35B Instruct?

Grok 4.1 Fast is cheaper. It lists at $0.20 per million input tokens and $0.50 per million output tokens; Qwen3-Coder 480B-A35B Instruct lists at $1.50 and $7.50.

Is Grok 4.1 Fast or Qwen3-Coder 480B-A35B Instruct better for coding?

Qwen3-Coder 480B-A35B Instruct scores higher on coding benchmarks: 35.5 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Qwen3-Coder 480B-A35B Instruct does, with 262K tokens against 128K.

How many benchmarks do Grok 4.1 Fast and Qwen3-Coder 480B-A35B Instruct share?

19 benchmarks have published results for both models. Grok 4.1 Fast has 32 scored results on Noometry and Qwen3-Coder 480B-A35B Instruct has 25.

Related comparisons

Go deeper