Model comparison

Grok Build 0.1 vs Wizardlm 70b

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 33.0 on the Noometry Index.

Last verified . 0 shared benchmarks.

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • The widest gap is in coding, where Grok Build 0.1 leads 43.1 to 31.4.
  • Wizardlm 70b has downloadable open weights; the other is API-only.

Side by side

Grok Build 0.1 and Wizardlm 70b specifications
Grok Build 0.1Wizardlm 70b
ProviderxAIMicrosoft
Noometry Index36.433.0
Released2026-04-16—
WeightsProprietaryOpen
Context window256K—
Max output256K—
Input $ / M tokens$1—
Output $ / M tokens$2—
Results tracked312

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Grok Build 0.1: 43.1 (#91), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkGrok Build 0.1Wizardlm 70b
SciCode50.2%—
LMArena Coding—1081

Agentic & Tool Use Not comparable

Grok Build 0.1: 22.7 (#129), Wizardlm 70b: —

Agentic & Tool Use benchmarks
BenchmarkGrok Build 0.1Wizardlm 70b
GBAEval2.4%—

Reasoning Grok Build 0.1 leads

Grok Build 0.1: 32.2 (#77), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkGrok Build 0.1Wizardlm 70b
CritPt9.1%—
LMArena Hard Prompts—1079

Math Not comparable

Grok Build 0.1: —, Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkGrok Build 0.1Wizardlm 70b
LMArena Math—1116

Multilingual Not comparable

Grok Build 0.1: —, Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkGrok Build 0.1Wizardlm 70b
LMArena Non-English—1078
LMArena Chinese—1052
LMArena German—1083
LMArena Russian—1155

Instruction Following Not comparable

Grok Build 0.1: —, Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkGrok Build 0.1Wizardlm 70b
LMArena Instruction Following—1093

Long Context Not comparable

Grok Build 0.1: —, Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkGrok Build 0.1Wizardlm 70b
LMArena Longer Query—1097

Writing & Preference Not comparable

Grok Build 0.1: —, Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkGrok Build 0.1Wizardlm 70b
LMArena Text—1120
LMArena Creative Writing—1149
LMArena Multi-Turn—1108

Frequently asked questions

Is Grok Build 0.1 better than Wizardlm 70b?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 33.0 on the Noometry Index.

Is Grok Build 0.1 or Wizardlm 70b better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 31.4 in the Noometry coding category.

How many benchmarks do Grok Build 0.1 and Wizardlm 70b share?

0 benchmarks have published results for both models. Grok Build 0.1 has 3 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper