Model comparison

Grok 4.1 vs Qwen3.5 27B

Grok 4.1 and Qwen3.5 27B score almost the same on the Noometry Index (41.5 vs 41.9), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Qwen3.5 27B Alibaba (Qwen)

41.9

Rank #127 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok 4.1 scores higher in 7 categories and Qwen3.5 27B in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Qwen3.5 27B leads 38.9 to 33.7.
  • Qwen3.5 27B has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 and Qwen3.5 27B specifications
Grok 4.1Qwen3.5 27B
ProviderxAIAlibaba (Qwen)
Noometry Index41.541.9
Released2025-11-172026-02-23
WeightsProprietaryOpen
Context window—262K
Max output—66K
Input $ / M tokens—$0.30
Output $ / M tokens—$2.40
Results tracked1928

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 27B leads

Grok 4.1: 33.7 (#253), Qwen3.5 27B: 38.9 (#168)

Coding benchmarks
BenchmarkGrok 4.1Qwen3.5 27B
LMArena WebDev12141358
LMArena Coding14451427
WeirdML—39.5%
ALE-Bench—349.45

Agentic & Tool Use Not comparable

Grok 4.1: 34.1 (#49), Qwen3.5 27B: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Qwen3.5 27B
Cybench39%—
Vending-Bench 2—201.98

Reasoning Grok 4.1 leads

Grok 4.1: 29.5 (#91), Qwen3.5 27B: 27.5 (#117)

Reasoning benchmarks
BenchmarkGrok 4.1Qwen3.5 27B
LMArena Hard Prompts14351414
NYT Connections (extended)—47.9%
Thematic Generalization—45.5%
DTBench—82.4%
LMCA—34%

Math Too close to call

Grok 4.1: 38.9 (#120), Qwen3.5 27B: 38.8 (#127)

Math benchmarks
BenchmarkGrok 4.1Qwen3.5 27B
LMArena Math14221429
MathArena Final-Answer Competitions—56.7%

Knowledge Grok 4.1 leads

Grok 4.1: 39.5 (#133), Qwen3.5 27B: 38.0 (#150)

Knowledge benchmarks
BenchmarkGrok 4.1Qwen3.5 27B
LMArena Expert14171428
Vectara Hallucination Rate—12.1%

Multimodal Not comparable

Grok 4.1: —, Qwen3.5 27B: 39.4 (#59)

Multimodal benchmarks
BenchmarkGrok 4.1Qwen3.5 27B
LMArena Vision—1241

Multilingual Grok 4.1 leads

Grok 4.1: 53.4 (#68), Qwen3.5 27B: 50.8 (#115)

Multilingual benchmarks
BenchmarkGrok 4.1Qwen3.5 27B
LMArena Non-English14251390
LMArena Chinese14651478
LMArena French14481410
LMArena German14461393
LMArena Japanese13971345
LMArena Korean14071358
LMArena Russian14341390
LMArena Spanish14381407

Instruction Following Too close to call

Grok 4.1: 73.8 (#111), Qwen3.5 27B: 73.5 (#119)

Instruction Following benchmarks
BenchmarkGrok 4.1Qwen3.5 27B
LMArena Instruction Following14001393

Long Context Too close to call

Grok 4.1: 43.2 (#100), Qwen3.5 27B: 43.1 (#106)

Long Context benchmarks
BenchmarkGrok 4.1Qwen3.5 27B
LMArena Longer Query14161413

Writing & Preference Grok 4.1 leads

Grok 4.1: 62.4 (#75), Qwen3.5 27B: 59.3 (#111)

Writing & Preference benchmarks
BenchmarkGrok 4.1Qwen3.5 27B
LMArena Text14371409
LMArena Creative Writing14111362
LMArena Multi-Turn14371410

Frequently asked questions

Is Grok 4.1 better than Qwen3.5 27B?

Grok 4.1 and Qwen3.5 27B score almost the same on the Noometry Index (41.5 vs 41.9), so choose on price, context window or the category you care about most.

Is Grok 4.1 or Qwen3.5 27B better for coding?

Qwen3.5 27B scores higher on coding benchmarks: 38.9 versus 33.7 in the Noometry coding category.

How many benchmarks do Grok 4.1 and Qwen3.5 27B share?

18 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Qwen3.5 27B has 28.

Related comparisons

Go deeper