Model comparison

Inkling vs Llama 13b

Inkling is the stronger model overall, scoring 44.1 to 24.4 on the Noometry Index.

Last verified . 9 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Inkling scores higher in 6 categories and Llama 13b in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Inkling leads 65.2 to 13.8.

Side by side

Inkling and Llama 13b specifications
InklingLlama 13b
ProviderThinking Machines LabMeta
Noometry Index44.124.4
Released2026-07-152023-02-24
WeightsOpenOpen
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked4121

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling leads

Inkling: 34.5 (#234), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkInklingLlama 13b
LMArena Coding1464683
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Llama 13b: —

Agentic & Tool Use benchmarks
BenchmarkInklingLlama 13b
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkInklingLlama 13b
LMArena Hard Prompts1451728
Epoch Capabilities Index148.54100.58
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
BIG-Bench Hard—37.9%
HellaSwag—79.2%
LAMBADA—75.2%
PIQA—80.1%
WinoGrande—73%

Math Inkling leads

Inkling: 31.3 (#225), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkInklingLlama 13b
LMArena Math1479838
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—
GSM8K—20.6%

Knowledge Not comparable

Inkling: 55.1 (#49), Llama 13b: —

Knowledge benchmarks
BenchmarkInklingLlama 13b
GPQA Diamond88.3%—
SimpleQA Verified40.3%—
LMArena Expert1465—
ARC (AI2) Challenge—52.7%
BoolQ—78.7%
MMLU—47.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

Inkling: —, Llama 13b: —

Multimodal benchmarks
BenchmarkInklingLlama 13b
ScienceQA—43.3%

Multilingual Inkling leads

Inkling: 54.0 (#52), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkInklingLlama 13b
LMArena Non-English1434819
LMArena Chinese1490—
LMArena French1458—
LMArena German1446—
LMArena Japanese1429—
LMArena Korean1404—
LMArena Russian1429—
LMArena Spanish1448—

Instruction Following Inkling leads

Inkling: 75.1 (#71), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkInklingLlama 13b
LMArena Instruction Following1426781

Long Context Not comparable

Inkling: 43.8 (#86), Llama 13b: —

Long Context benchmarks
BenchmarkInklingLlama 13b
LMArena Longer Query1434—

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkInklingLlama 13b
LMArena Text1441834
LMArena Creative Writing1387794
LMArena Multi-Turn1436753
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Llama 13b?

Inkling is the stronger model overall, scoring 44.1 to 24.4 on the Noometry Index.

Is Inkling or Llama 13b better for coding?

Inkling scores higher on coding benchmarks: 34.5 versus 21.4 in the Noometry coding category.

How many benchmarks do Inkling and Llama 13b share?

9 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper