Model comparison

Grok 4.1 vs Qwen3 32B

Grok 4.1 is the stronger model overall, scoring 41.5 to 39.2 on the Noometry Index.

Last verified . 13 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Grok 4.1 scores higher in 5 categories and Qwen3 32B in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Grok 4.1 leads 62.4 to 52.9.
  • Qwen3 32B has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 and Qwen3 32B specifications
Grok 4.1Qwen3 32B
ProviderxAIAlibaba (Qwen)
Noometry Index41.539.2
Released2025-11-172025-04
WeightsProprietaryOpen
Context window—131K
Max output—16K
Input $ / M tokens—$0.70
Output $ / M tokens—$2.80
Results tracked1926

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 32B leads

Grok 4.1: 33.7 (#253), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkGrok 4.1Qwen3 32B
LMArena Coding14451358
Aider Polyglot—40%
LMArena WebDev1214—
SciCode—35.4%

Agentic & Tool Use Grok 4.1 leads

Grok 4.1: 34.1 (#49), Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Qwen3 32B
Berkeley Function Calling Leaderboard—48.7%
Cybench39%—

Reasoning Grok 4.1 leads

Grok 4.1: 29.5 (#91), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkGrok 4.1Qwen3 32B
LMArena Hard Prompts14351334
Kagi LLM Benchmark—54.9%
CritPt—0.3%
Chess Puzzles—5%
DTBench—67.5%
LMCA—17.3%
Epoch Capabilities Index—138.51

Math Too close to call

Grok 4.1: 38.9 (#120), Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkGrok 4.1Qwen3 32B
LMArena Math14221399
OTIS Mock AIME 2024-2025—66.9%

Knowledge Too close to call

Grok 4.1: 39.5 (#133), Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkGrok 4.1Qwen3 32B
LMArena Expert14171362
GPQA Diamond—65.7%
Vectara Hallucination Rate—5.9%

Multilingual Grok 4.1 leads

Grok 4.1: 53.4 (#68), Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkGrok 4.1Qwen3 32B
LMArena Non-English14251317
LMArena Chinese14651357
LMArena German14461341
LMArena Russian14341311
LMArena French1448—
LMArena Japanese1397—
LMArena Korean1407—
LMArena Spanish1438—

Instruction Following Grok 4.1 leads

Grok 4.1: 73.8 (#111), Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkGrok 4.1Qwen3 32B
LMArena Instruction Following14001305

Long Context Too close to call

Grok 4.1: 43.2 (#100), Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkGrok 4.1Qwen3 32B
LMArena Longer Query14161327
Fiction.LiveBench—74.2%

Writing & Preference Grok 4.1 leads

Grok 4.1: 62.4 (#75), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkGrok 4.1Qwen3 32B
LMArena Text14371340
LMArena Creative Writing14111297
LMArena Multi-Turn14371331

Frequently asked questions

Is Grok 4.1 better than Qwen3 32B?

Grok 4.1 is the stronger model overall, scoring 41.5 to 39.2 on the Noometry Index.

Is Grok 4.1 or Qwen3 32B better for coding?

Qwen3 32B scores higher on coding benchmarks: 37.7 versus 33.7 in the Noometry coding category.

How many benchmarks do Grok 4.1 and Qwen3 32B share?

13 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper