Model comparison

Grok 4.3 vs Wizardlm 13b

Grok 4.3 is the stronger model overall, scoring 43.8 to 31.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Grok 4.3 scores higher in 7 categories and Wizardlm 13b in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Grok 4.3 leads 58.5 to 30.5.
  • Wizardlm 13b has downloadable open weights; the other is API-only.

Side by side

Grok 4.3 and Wizardlm 13b specifications
Grok 4.3Wizardlm 13b
ProviderxAIMicrosoft
Noometry Index43.831.4
Released2026-04-17—
WeightsProprietaryOpen
Context window1M—
Max output30K—
Input $ / M tokens$1.25—
Output $ / M tokens$2.50—
Results tracked4010

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

Grok 4.3: 41.6 (#121), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkGrok 4.3Wizardlm 13b
LMArena Coding14151035
LMArena WebDev1357—
SciCode47.3%—
WeirdML49.9%—
ALE-Bench944.17—

Agentic & Tool Use Not comparable

Grok 4.3: 27.7 (#99), Wizardlm 13b: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Wizardlm 13b
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkGrok 4.3Wizardlm 13b
LMArena Hard Prompts13961018
NYT Connections (extended)55.2%—
CritPt8%—
Chess Puzzles25%—
DTBench90.7%—
LMCA38.3%—
Epoch Capabilities Index149.16—
ForecastBench60.3—

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkGrok 4.3Wizardlm 13b
LMArena Math13881017
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
OTIS Mock AIME 2024-202593.3%—
ProofBench11%—

Knowledge Not comparable

Grok 4.3: 52.5 (#62), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkGrok 4.3Wizardlm 13b
GPQA Diamond88.8%—
SimpleQA Verified33.2%—
LMArena Expert1385—

Multimodal Not comparable

Grok 4.3: 31.6 (#104), Wizardlm 13b: —

Multimodal benchmarks
BenchmarkGrok 4.3Wizardlm 13b
LMArena Vision1229—
Blueprint-Bench 20%—

Multilingual Grok 4.3 leads

Grok 4.3: 50.5 (#120), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkGrok 4.3Wizardlm 13b
LMArena Non-English13851034
LMArena Chinese14221023
LMArena French1412—
LMArena German1395—
LMArena Japanese1379—
LMArena Korean1356—
LMArena Russian1399—
LMArena Spanish1398—

Instruction Following Grok 4.3 leads

Grok 4.3: 72.1 (#140), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkGrok 4.3Wizardlm 13b
LMArena Instruction Following13661048

Long Context Grok 4.3 leads

Grok 4.3: 42.5 (#123), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkGrok 4.3Wizardlm 13b
LMArena Longer Query13931054

Writing & Preference Grok 4.3 leads

Grok 4.3: 58.5 (#118), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkGrok 4.3Wizardlm 13b
LMArena Text13971077
LMArena Creative Writing13801091
LMArena Multi-Turn14061047
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Wizardlm 13b?

Grok 4.3 is the stronger model overall, scoring 43.8 to 31.4 on the Noometry Index.

Is Grok 4.3 or Wizardlm 13b better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 30.1 in the Noometry coding category.

How many benchmarks do Grok 4.3 and Wizardlm 13b share?

10 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper