Model comparison

Grok 2 Mini 2024 08 13 vs Qwen Max

Grok 2 Mini 2024 08 13 is the stronger model overall, scoring 37.7 to 34.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Grok 2 Mini 2024 08 13 xAI

37.7

Rank #198 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Grok 2 Mini 2024 08 13 scores higher in 3 categories and Qwen Max in 5 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 2 Mini 2024 08 13 leads 35.4 to 22.3.

Side by side

Grok 2 Mini 2024 08 13 and Qwen Max specifications
Grok 2 Mini 2024 08 13Qwen Max
ProviderxAIAlibaba (Qwen)
Noometry Index37.734.7
Released2024-08-132024-04-03
WeightsProprietaryProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked1723

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 37.0 (#199), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkGrok 2 Mini 2024 08 13Qwen Max
LMArena Coding12691288
Aider Polyglot—21.8%

Reasoning Too close to call

Grok 2 Mini 2024 08 13: 24.8 (#159), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkGrok 2 Mini 2024 08 13Qwen Max
LMArena Hard Prompts12551269

Math Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 35.4 (#185), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkGrok 2 Mini 2024 08 13Qwen Max
LMArena Math12651275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Grok 2 Mini 2024 08 13 leads

Grok 2 Mini 2024 08 13: 33.9 (#200), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkGrok 2 Mini 2024 08 13Qwen Max
LMArena Expert12381248
GPQA Diamond—56.1%

Multilingual Too close to call

Grok 2 Mini 2024 08 13: 41.4 (#208), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkGrok 2 Mini 2024 08 13Qwen Max
LMArena Non-English12571263
LMArena Chinese12621254
LMArena French12861330
LMArena German12731254
LMArena Japanese12131205
LMArena Korean11951142
LMArena Russian12611274
LMArena Spanish12791290

Instruction Following Too close to call

Grok 2 Mini 2024 08 13: 65.5 (#220), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkGrok 2 Mini 2024 08 13Qwen Max
LMArena Instruction Following12451262

Long Context Too close to call

Grok 2 Mini 2024 08 13: 38.4 (#196), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkGrok 2 Mini 2024 08 13Qwen Max
LMArena Longer Query12661288
Fiction.LiveBench—66.7%

Writing & Preference Too close to call

Grok 2 Mini 2024 08 13: 47.3 (#211), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkGrok 2 Mini 2024 08 13Qwen Max
LMArena Text12811282
LMArena Creative Writing12421248
LMArena Multi-Turn12651277

Frequently asked questions

Is Grok 2 Mini 2024 08 13 better than Qwen Max?

Grok 2 Mini 2024 08 13 is the stronger model overall, scoring 37.7 to 34.7 on the Noometry Index.

Is Grok 2 Mini 2024 08 13 or Qwen Max better for coding?

Grok 2 Mini 2024 08 13 scores higher on coding benchmarks: 37.0 versus 30.7 in the Noometry coding category.

How many benchmarks do Grok 2 Mini 2024 08 13 and Qwen Max share?

17 benchmarks have published results for both models. Grok 2 Mini 2024 08 13 has 17 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper