Model comparison

Grok Build 0.1 vs Qwen3 235B-A22B

Qwen3 235B-A22B is the stronger model overall, scoring 43.5 to 36.4 on the Noometry Index.

Last verified . 2 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Qwen3 235B-A22B Alibaba (Qwen)

43.5

Rank #91 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Grok Build 0.1 scores higher in 1 category and Qwen3 235B-A22B in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 15.7.
  • The biggest single-benchmark swing is CritPt: 9.1% for Grok Build 0.1 and 0% for Qwen3 235B-A22B.
  • Both cost about the same: $1 input and $2 output per million tokens.
  • Grok Build 0.1 accepts more context: 256K tokens versus 131K.
  • Qwen3 235B-A22B has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Qwen3 235B-A22B specifications
Grok Build 0.1Qwen3 235B-A22B
ProviderxAIAlibaba (Qwen)
Noometry Index36.443.5
Released2026-04-162025-04
WeightsProprietaryOpen
Context window256K131K
Max output256K16K
Input $ / M tokens$1$0.70
Output $ / M tokens$2$2.80
Results tracked349

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 235B-A22B leads

Grok Build 0.1: 43.1 (#91), Qwen3 235B-A22B: 44.3 (#75)

Coding benchmarks
BenchmarkGrok Build 0.1Qwen3 235B-A22B
SciCode50.2%42.4%
Aider Polyglot—59.6%
WeirdML—41%
LMArena Coding—1445

Agentic & Tool Use Qwen3 235B-A22B leads

Grok Build 0.1: 22.7 (#129), Qwen3 235B-A22B: 33.9 (#51)

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Qwen3 235B-A22B
Berkeley Function Calling Leaderboard—52.1%
GBAEval2.4%—
Vending-Bench 2—-11.34

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Qwen3 235B-A22B: 15.7 (#311)

Reasoning benchmarks
BenchmarkGrok Build 0.1Qwen3 235B-A22B
CritPt9.1%0%
ARC-AGI-2—1.3%
SimpleBench—31%
Kagi LLM Benchmark—69.4%
ARC-AGI-1—11%
Chess Puzzles—12%
LMArena Hard Prompts—1433
Mystery Game Puzzles—9%
DTBench—80.3%
LMCA—29.3%
Epoch Capabilities Index—143.85
ForecastBench—59.7

Math Not comparable

Grok Build 0.1: —, Qwen3 235B-A22B: 50.4 (#57)

Math benchmarks
BenchmarkGrok Build 0.1Qwen3 235B-A22B
OTIS Mock AIME 2024-2025—86.7%
Omni-MATH—71.8%
LMArena Math—1432
MATH Level 5—68.9%
FrontierMath (Feb 2025 set)—8.5%
FrontierMath Tier 4 (v1)—0%

Knowledge Not comparable

Grok Build 0.1: —, Qwen3 235B-A22B: 49.6 (#73)

Knowledge benchmarks
BenchmarkGrok Build 0.1Qwen3 235B-A22B
GPQA Diamond—80.1%
SimpleQA Verified—40.4%
MMLU-Pro—84.4%
Confabulations—15.6%
Vectara Hallucination Rate—9.3%
GPQA (HELM)—72.7%
LMArena Expert—1463

Multilingual Not comparable

Grok Build 0.1: —, Qwen3 235B-A22B: 52.3 (#89)

Multilingual benchmarks
BenchmarkGrok Build 0.1Qwen3 235B-A22B
LMArena Non-English—1409
LMArena Chinese—1481
LMArena French—1445
LMArena German—1433
LMArena Japanese—1399
LMArena Korean—1391
LMArena Russian—1411
LMArena Spanish—1430

Instruction Following Not comparable

Grok Build 0.1: —, Qwen3 235B-A22B: 72.6 (#136)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Qwen3 235B-A22B
IFEval—83.5%
LMArena Instruction Following—1408

Long Context Not comparable

Grok Build 0.1: —, Qwen3 235B-A22B: 46.1 (#26)

Long Context benchmarks
BenchmarkGrok Build 0.1Qwen3 235B-A22B
Fiction.LiveBench—75%
LMArena Longer Query—1426

Writing & Preference Not comparable

Grok Build 0.1: —, Qwen3 235B-A22B: 59.6 (#108)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Qwen3 235B-A22B
LMArena Text—1419
LMArena Creative Writing—1384
Short-Story Creative Writing—83%
EQ-Bench Creative Writing—1366
WildBench—86.6%
LMArena Multi-Turn—1432

Frequently asked questions

Is Grok Build 0.1 better than Qwen3 235B-A22B?

Qwen3 235B-A22B is the stronger model overall, scoring 43.5 to 36.4 on the Noometry Index.

Which is cheaper, Grok Build 0.1 or Qwen3 235B-A22B?

Qwen3 235B-A22B is cheaper. It lists at $0.70 per million input tokens and $2.80 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Grok Build 0.1 or Qwen3 235B-A22B better for coding?

Qwen3 235B-A22B scores higher on coding benchmarks: 44.3 versus 43.1 in the Noometry coding category.

Which has the bigger context window?

Grok Build 0.1 does, with 256K tokens against 131K.

How many benchmarks do Grok Build 0.1 and Qwen3 235B-A22B share?

2 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Qwen3 235B-A22B has 49.

Related comparisons

Go deeper