Model comparison
Grok 4.1 vs Qwen3 32B
Grok 4.1 is the stronger model overall, scoring 41.5 to 39.2 on the Noometry Index.
Last verified . 13 shared benchmarks.
Summary
- They share 13 benchmarks with published results for both. Grok 4.1 scores higher in 5 categories and Qwen3 32B in 4 categories; 6 gaps are clear of the uncertainty.
- The widest gap is in writing & preference, where Grok 4.1 leads 62.4 to 52.9.
- Qwen3 32B has downloadable open weights; the other is API-only.
Side by side
| Grok 4.1 | Qwen3 32B | |
|---|---|---|
| Provider | xAI | Alibaba (Qwen) |
| Noometry Index | 41.5 | 39.2 |
| Released | 2025-11-17 | 2025-04 |
| Weights | Proprietary | Open |
| Context window | — | 131K |
| Max output | — | 16K |
| Input $ / M tokens | — | $0.70 |
| Output $ / M tokens | — | $2.80 |
| Results tracked | 19 | 26 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Qwen3 32B leads
Grok 4.1: 33.7 (#253), Qwen3 32B: 37.7 (#190)
| Benchmark | Grok 4.1 | Qwen3 32B |
|---|---|---|
| LMArena Coding | 1445 | 1358 |
| Aider Polyglot | — | 40% |
| LMArena WebDev | 1214 | — |
| SciCode | — | 35.4% |
Agentic & Tool Use Grok 4.1 leads
Grok 4.1: 34.1 (#49), Qwen3 32B: 32.6 (#62)
| Benchmark | Grok 4.1 | Qwen3 32B |
|---|---|---|
| Berkeley Function Calling Leaderboard | — | 48.7% |
| Cybench | 39% | — |
Reasoning Grok 4.1 leads
Grok 4.1: 29.5 (#91), Qwen3 32B: 20.2 (#241)
| Benchmark | Grok 4.1 | Qwen3 32B |
|---|---|---|
| LMArena Hard Prompts | 1435 | 1334 |
| Kagi LLM Benchmark | — | 54.9% |
| CritPt | — | 0.3% |
| Chess Puzzles | — | 5% |
| DTBench | — | 67.5% |
| LMCA | — | 17.3% |
| Epoch Capabilities Index | — | 138.51 |
Math Too close to call
Grok 4.1: 38.9 (#120), Qwen3 32B: 39.7 (#99)
| Benchmark | Grok 4.1 | Qwen3 32B |
|---|---|---|
| LMArena Math | 1422 | 1399 |
| OTIS Mock AIME 2024-2025 | — | 66.9% |
Knowledge Too close to call
Grok 4.1: 39.5 (#133), Qwen3 32B: 40.0 (#125)
| Benchmark | Grok 4.1 | Qwen3 32B |
|---|---|---|
| LMArena Expert | 1417 | 1362 |
| GPQA Diamond | — | 65.7% |
| Vectara Hallucination Rate | — | 5.9% |
Multilingual Grok 4.1 leads
Grok 4.1: 53.4 (#68), Qwen3 32B: 45.6 (#167)
| Benchmark | Grok 4.1 | Qwen3 32B |
|---|---|---|
| LMArena Non-English | 1425 | 1317 |
| LMArena Chinese | 1465 | 1357 |
| LMArena German | 1446 | 1341 |
| LMArena Russian | 1434 | 1311 |
| LMArena French | 1448 | — |
| LMArena Japanese | 1397 | — |
| LMArena Korean | 1407 | — |
| LMArena Spanish | 1438 | — |
Instruction Following Grok 4.1 leads
Grok 4.1: 73.8 (#111), Qwen3 32B: 68.9 (#179)
| Benchmark | Grok 4.1 | Qwen3 32B |
|---|---|---|
| LMArena Instruction Following | 1400 | 1305 |
Long Context Too close to call
Grok 4.1: 43.2 (#100), Qwen3 32B: 43.8 (#87)
| Benchmark | Grok 4.1 | Qwen3 32B |
|---|---|---|
| LMArena Longer Query | 1416 | 1327 |
| Fiction.LiveBench | — | 74.2% |
Writing & Preference Grok 4.1 leads
Grok 4.1: 62.4 (#75), Qwen3 32B: 52.9 (#163)
| Benchmark | Grok 4.1 | Qwen3 32B |
|---|---|---|
| LMArena Text | 1437 | 1340 |
| LMArena Creative Writing | 1411 | 1297 |
| LMArena Multi-Turn | 1437 | 1331 |
Frequently asked questions
Is Grok 4.1 better than Qwen3 32B?
Grok 4.1 is the stronger model overall, scoring 41.5 to 39.2 on the Noometry Index.
Is Grok 4.1 or Qwen3 32B better for coding?
Qwen3 32B scores higher on coding benchmarks: 37.7 versus 33.7 in the Noometry coding category.
How many benchmarks do Grok 4.1 and Qwen3 32B share?
13 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Qwen3 32B has 26.