Model comparison

Grok 4.1 vs Wizardlm 70b

Grok 4.1 is the stronger model overall, scoring 41.5 to 33.0 on the Noometry Index.

Last verified . 12 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Grok 4.1 scores higher in 7 categories and Wizardlm 70b in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Grok 4.1 leads 62.4 to 34.8.
  • Wizardlm 70b has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 and Wizardlm 70b specifications
Grok 4.1Wizardlm 70b
ProviderxAIMicrosoft
Noometry Index41.533.0
Released2025-11-17—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1912

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.1 leads

Grok 4.1: 33.7 (#253), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkGrok 4.1Wizardlm 70b
LMArena Coding14451081
LMArena WebDev1214—

Agentic & Tool Use Not comparable

Grok 4.1: 34.1 (#49), Wizardlm 70b: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Wizardlm 70b
Cybench39%—

Reasoning Grok 4.1 leads

Grok 4.1: 29.5 (#91), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkGrok 4.1Wizardlm 70b
LMArena Hard Prompts14351079

Math Grok 4.1 leads

Grok 4.1: 38.9 (#120), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkGrok 4.1Wizardlm 70b
LMArena Math14221116

Knowledge Not comparable

Grok 4.1: 39.5 (#133), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkGrok 4.1Wizardlm 70b
LMArena Expert1417—

Multilingual Grok 4.1 leads

Grok 4.1: 53.4 (#68), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkGrok 4.1Wizardlm 70b
LMArena Non-English14251078
LMArena Chinese14651052
LMArena German14461083
LMArena Russian14341155
LMArena French1448—
LMArena Japanese1397—
LMArena Korean1407—
LMArena Spanish1438—

Instruction Following Grok 4.1 leads

Grok 4.1: 73.8 (#111), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkGrok 4.1Wizardlm 70b
LMArena Instruction Following14001093

Long Context Grok 4.1 leads

Grok 4.1: 43.2 (#100), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkGrok 4.1Wizardlm 70b
LMArena Longer Query14161097

Writing & Preference Grok 4.1 leads

Grok 4.1: 62.4 (#75), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkGrok 4.1Wizardlm 70b
LMArena Text14371120
LMArena Creative Writing14111149
LMArena Multi-Turn14371108

Frequently asked questions

Is Grok 4.1 better than Wizardlm 70b?

Grok 4.1 is the stronger model overall, scoring 41.5 to 33.0 on the Noometry Index.

Is Grok 4.1 or Wizardlm 70b better for coding?

Grok 4.1 scores higher on coding benchmarks: 33.7 versus 31.4 in the Noometry coding category.

How many benchmarks do Grok 4.1 and Wizardlm 70b share?

12 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper