Model comparison

Inkling vs Llama 3.1 Nemotron 51b Instruct

Inkling is the stronger model overall, scoring 44.1 to 35.9 on the Noometry Index.

Last verified . 12 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Inkling scores higher in 6 categories and Llama 3.1 Nemotron 51b Instruct in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 31.9.

Side by side

Inkling and Llama 3.1 Nemotron 51b Instruct specifications
InklingLlama 3.1 Nemotron 51b Instruct
ProviderThinking Machines LabNVIDIA
Noometry Index44.135.9
Released2026-07-15—
WeightsOpenOpen
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked4112

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Nemotron 51b Instruct leads

Inkling: 34.5 (#234), Llama 3.1 Nemotron 51b Instruct: 35.6 (#222)

Coding benchmarks
BenchmarkInklingLlama 3.1 Nemotron 51b Instruct
LMArena Coding14641223
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Llama 3.1 Nemotron 51b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkInklingLlama 3.1 Nemotron 51b Instruct
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Llama 3.1 Nemotron 51b Instruct: 23.5 (#177)

Reasoning benchmarks
BenchmarkInklingLlama 3.1 Nemotron 51b Instruct
LMArena Hard Prompts14511203
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Llama 3.1 Nemotron 51b Instruct leads

Inkling: 31.3 (#225), Llama 3.1 Nemotron 51b Instruct: 34.6 (#193)

Math benchmarks
BenchmarkInklingLlama 3.1 Nemotron 51b Instruct
LMArena Math14791230
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Llama 3.1 Nemotron 51b Instruct: 31.9 (#218)

Knowledge benchmarks
BenchmarkInklingLlama 3.1 Nemotron 51b Instruct
LMArena Expert14651167
GPQA Diamond88.3%—
SimpleQA Verified40.3%—

Multilingual Inkling leads

Inkling: 54.0 (#52), Llama 3.1 Nemotron 51b Instruct: 36.1 (#241)

Multilingual benchmarks
BenchmarkInklingLlama 3.1 Nemotron 51b Instruct
LMArena Non-English14341181
LMArena Chinese14901180
LMArena Russian14291187
LMArena French1458—
LMArena German1446—
LMArena Japanese1429—
LMArena Korean1404—
LMArena Spanish1448—

Instruction Following Inkling leads

Inkling: 75.1 (#71), Llama 3.1 Nemotron 51b Instruct: 62.9 (#233)

Instruction Following benchmarks
BenchmarkInklingLlama 3.1 Nemotron 51b Instruct
LMArena Instruction Following14261201

Long Context Inkling leads

Inkling: 43.8 (#86), Llama 3.1 Nemotron 51b Instruct: 36.5 (#230)

Long Context benchmarks
BenchmarkInklingLlama 3.1 Nemotron 51b Instruct
LMArena Longer Query14341205

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Llama 3.1 Nemotron 51b Instruct: 43.4 (#229)

Writing & Preference benchmarks
BenchmarkInklingLlama 3.1 Nemotron 51b Instruct
LMArena Text14411228
LMArena Creative Writing13871213
LMArena Multi-Turn14361227
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Llama 3.1 Nemotron 51b Instruct?

Inkling is the stronger model overall, scoring 44.1 to 35.9 on the Noometry Index.

Is Inkling or Llama 3.1 Nemotron 51b Instruct better for coding?

Llama 3.1 Nemotron 51b Instruct scores higher on coding benchmarks: 35.6 versus 34.5 in the Noometry coding category.

How many benchmarks do Inkling and Llama 3.1 Nemotron 51b Instruct share?

12 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Llama 3.1 Nemotron 51b Instruct has 12.

Related comparisons

Go deeper