Model comparison

Grok-3 mini vs Qwen Plus

Grok-3 mini is the stronger model overall, scoring 41.2 to 37.1 on the Noometry Index.

Last verified . 18 shared benchmarks.

Grok-3 mini xAI

41.2

Rank #141 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok-3 mini scores higher in 7 categories and Qwen Plus in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok-3 mini leads 46.4 to 27.4.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 77.8% for Grok-3 mini and 17.8% for Qwen Plus.

Side by side

Grok-3 mini and Qwen Plus specifications
Grok-3 miniQwen Plus
ProviderxAIAlibaba (Qwen)
Noometry Index41.237.1
Released2025-04-092024-01-25
WeightsProprietaryProprietary
Context window—1M
Max output—33K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.20
Results tracked3520

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok-3 mini leads

Grok-3 mini: 40.8 (#131), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkGrok-3 miniQwen Plus
LMArena Coding13791328
Aider Polyglot49.3%—
WeirdML42.6%—

Reasoning Qwen Plus leads

Grok-3 mini: 13.6 (#334), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkGrok-3 miniQwen Plus
Kagi LLM Benchmark61.3%63.3%
LMArena Hard Prompts13751317
ARC-AGI-20.4%—
ARC-AGI-116.5%—
DTBench—81.1%
LMCA—24%
Epoch Capabilities Index140.35—

Math Grok-3 mini leads

Grok-3 mini: 42.1 (#85), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkGrok-3 miniQwen Plus
OTIS Mock AIME 2024-202577.8%17.8%
LMArena Math13861326
MATH Level 590.9%65.3%
FrontierMath (Feb 2025 set)5.9%1.7%
Omni-MATH31.8%—

Knowledge Grok-3 mini leads

Grok-3 mini: 46.4 (#81), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkGrok-3 miniQwen Plus
GPQA Diamond76.3%48.1%
LMArena Expert13951328
MMLU-Pro79.9%—
Confabulations10.8%—
GPQA (HELM)67.5%—

Multilingual Grok-3 mini leads

Grok-3 mini: 48.1 (#145), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkGrok-3 miniQwen Plus
LMArena Non-English13521310
LMArena Chinese13871347
LMArena Japanese13421251
LMArena Russian13531323
LMArena French1357—
LMArena German1349—
LMArena Korean1335—
LMArena Spanish1381—

Instruction Following Grok-3 mini leads

Grok-3 mini: 78.5 (#9), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkGrok-3 miniQwen Plus
LMArena Instruction Following13571303
IFEval95.1%—

Long Context Too close to call

Grok-3 mini: 41.0 (#147), Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkGrok-3 miniQwen Plus
LMArena Longer Query13721324
Fiction.LiveBench66.7%—

Writing & Preference Too close to call

Grok-3 mini: 52.5 (#169), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkGrok-3 miniQwen Plus
LMArena Text13701326
LMArena Creative Writing13421293
LMArena Multi-Turn13551336
Short-Story Creative Writing73.5%—
WildBench65.1%—

Frequently asked questions

Is Grok-3 mini better than Qwen Plus?

Grok-3 mini is the stronger model overall, scoring 41.2 to 37.1 on the Noometry Index.

Is Grok-3 mini or Qwen Plus better for coding?

Grok-3 mini scores higher on coding benchmarks: 40.8 versus 38.9 in the Noometry coding category.

How many benchmarks do Grok-3 mini and Qwen Plus share?

18 benchmarks have published results for both models. Grok-3 mini has 35 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper