Model comparison

Inkling vs Wizardlm 70b

Inkling is the stronger model overall, scoring 44.1 to 33.0 on the Noometry Index.

Last verified . 12 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Inkling scores higher in 6 categories and Wizardlm 70b in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Inkling leads 65.2 to 34.8.

Side by side

Inkling and Wizardlm 70b specifications
InklingWizardlm 70b
ProviderThinking Machines LabMicrosoft
Noometry Index44.133.0
Released2026-07-15—
WeightsOpenOpen
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked4112

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling leads

Inkling: 34.5 (#234), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkInklingWizardlm 70b
LMArena Coding14641081
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Wizardlm 70b: —

Agentic & Tool Use benchmarks
BenchmarkInklingWizardlm 70b
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkInklingWizardlm 70b
LMArena Hard Prompts14511079
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Too close to call

Inkling: 31.3 (#225), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkInklingWizardlm 70b
LMArena Math14791116
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—

Knowledge Not comparable

Inkling: 55.1 (#49), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkInklingWizardlm 70b
GPQA Diamond88.3%—
SimpleQA Verified40.3%—
LMArena Expert1465—

Multilingual Inkling leads

Inkling: 54.0 (#52), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkInklingWizardlm 70b
LMArena Non-English14341078
LMArena Chinese14901052
LMArena German14461083
LMArena Russian14291155
LMArena French1458—
LMArena Japanese1429—
LMArena Korean1404—
LMArena Spanish1448—

Instruction Following Inkling leads

Inkling: 75.1 (#71), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkInklingWizardlm 70b
LMArena Instruction Following14261093

Long Context Inkling leads

Inkling: 43.8 (#86), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkInklingWizardlm 70b
LMArena Longer Query14341097

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkInklingWizardlm 70b
LMArena Text14411120
LMArena Creative Writing13871149
LMArena Multi-Turn14361108
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Wizardlm 70b?

Inkling is the stronger model overall, scoring 44.1 to 33.0 on the Noometry Index.

Is Inkling or Wizardlm 70b better for coding?

Inkling scores higher on coding benchmarks: 34.5 versus 31.4 in the Noometry coding category.

How many benchmarks do Inkling and Wizardlm 70b share?

12 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper