Model comparison

Grok 4.1 vs Qwen3.6 Max Preview

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 41.5 on the Noometry Index.

Last verified . 15 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Grok 4.1 scores higher in 0 categories and Qwen3.6 Max Preview in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.6 Max Preview leads 57.6 to 39.5.

Side by side

Grok 4.1 and Qwen3.6 Max Preview specifications
Grok 4.1Qwen3.6 Max Preview
ProviderxAIAlibaba (Qwen)
Noometry Index41.551.5
Released2025-11-172026-04-20
WeightsProprietaryProprietary
Context window—262K
Max output—66K
Input $ / M tokens—$1.30
Output $ / M tokens—$7.80
Results tracked1929

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

Grok 4.1: 33.7 (#253), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkGrok 4.1Qwen3.6 Max Preview
LMArena WebDev12141482
LMArena Coding14451471
SWE-bench Verified—76.7%

Agentic & Tool Use Not comparable

Grok 4.1: 34.1 (#49), Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Qwen3.6 Max Preview
Cybench39%—
Vending-Bench 2—4,254

Reasoning Qwen3.6 Max Preview leads

Grok 4.1: 29.5 (#91), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkGrok 4.1Qwen3.6 Max Preview
LMArena Hard Prompts14351457
SimpleBench—63%
NYT Connections (extended)—74.1%
Chess Puzzles—20%
Mystery Game Puzzles—19%
DTBench—87.2%
LMCA—42.5%
Epoch Capabilities Index—149.24

Math Qwen3.6 Max Preview leads

Grok 4.1: 38.9 (#120), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkGrok 4.1Qwen3.6 Max Preview
LMArena Math14221465
OTIS Mock AIME 2024-2025—91.1%
FrontierMath (Feb 2025 set)—23.1%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Qwen3.6 Max Preview leads

Grok 4.1: 39.5 (#133), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkGrok 4.1Qwen3.6 Max Preview
LMArena Expert14171478
GPQA Diamond—87.4%
SimpleQA Verified—52%

Multilingual Too close to call

Grok 4.1: 53.4 (#68), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkGrok 4.1Qwen3.6 Max Preview
LMArena Non-English14251437
LMArena Chinese14651487
LMArena French14481449
LMArena Russian14341445
LMArena Spanish14381454
LMArena German1446—
LMArena Japanese1397—
LMArena Korean1407—

Instruction Following Qwen3.6 Max Preview leads

Grok 4.1: 73.8 (#111), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkGrok 4.1Qwen3.6 Max Preview
LMArena Instruction Following14001438

Long Context Qwen3.6 Max Preview leads

Grok 4.1: 43.2 (#100), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkGrok 4.1Qwen3.6 Max Preview
LMArena Longer Query14161457

Writing & Preference Qwen3.6 Max Preview leads

Grok 4.1: 62.4 (#75), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkGrok 4.1Qwen3.6 Max Preview
LMArena Text14371447
LMArena Creative Writing14111435
LMArena Multi-Turn14371456

Frequently asked questions

Is Grok 4.1 better than Qwen3.6 Max Preview?

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 41.5 on the Noometry Index.

Is Grok 4.1 or Qwen3.6 Max Preview better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 33.7 in the Noometry coding category.

How many benchmarks do Grok 4.1 and Qwen3.6 Max Preview share?

15 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper