Model comparison

Inkling vs Llama 3.2 3B

Inkling is the stronger model overall, scoring 44.1 to 28.9 on the Noometry Index. Llama 3.2 3B costs 21× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Last verified . 14 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Inkling scores higher in 8 categories and Llama 3.2 3B in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Inkling leads 65.2 to 24.7.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • Llama 3.2 3B accepts more context: 131K tokens versus 66K.

Side by side

Inkling and Llama 3.2 3B specifications
InklingLlama 3.2 3B
ProviderThinking Machines LabMeta
Noometry Index44.128.9
Released2026-07-152024-09-24
WeightsOpenOpen
Context window66K131K
Max output66K118K
Input $ / M tokens$1.87$0.05
Output $ / M tokens$4.68$0.33
Results tracked4118

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling leads

Inkling: 34.5 (#234), Llama 3.2 3B: 27.6 (#319)

Coding benchmarks
BenchmarkInklingLlama 3.2 3B
LMArena Coding14641098
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
BigCodeBench Instruct—23.4%
BigCodeBench Complete—28.3%
ALE-Bench946—

Agentic & Tool Use Inkling leads

Inkling: 29.6 (#85), Llama 3.2 3B: 20.1 (#143)

Agentic & Tool Use benchmarks
BenchmarkInklingLlama 3.2 3B
APEX-Agents33.8%—
Berkeley Function Calling Leaderboard—21.9%
τ²-bench Banking25%—
BALROG—10.1%

Reasoning Inkling leads

Inkling: 40.4 (#56), Llama 3.2 3B: 21.0 (#228)

Reasoning benchmarks
BenchmarkInklingLlama 3.2 3B
LMArena Hard Prompts14511095
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Llama 3.2 3B leads

Inkling: 31.3 (#225), Llama 3.2 3B: 32.4 (#214)

Math benchmarks
BenchmarkInklingLlama 3.2 3B
LMArena Math14791126
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Llama 3.2 3B: 29.7 (#235)

Knowledge benchmarks
BenchmarkInklingLlama 3.2 3B
LMArena Expert14651090
GPQA Diamond88.3%—
SimpleQA Verified40.3%—

Multilingual Inkling leads

Inkling: 54.0 (#52), Llama 3.2 3B: 26.2 (#281)

Multilingual benchmarks
BenchmarkInklingLlama 3.2 3B
LMArena Non-English14341019
LMArena Chinese14901017
LMArena German14461056
LMArena Russian1429949
LMArena French1458—
LMArena Japanese1429—
LMArena Korean1404—
LMArena Spanish1448—

Instruction Following Inkling leads

Inkling: 75.1 (#71), Llama 3.2 3B: 56.0 (#275)

Instruction Following benchmarks
BenchmarkInklingLlama 3.2 3B
LMArena Instruction Following14261089

Long Context Inkling leads

Inkling: 43.8 (#86), Llama 3.2 3B: 33.4 (#261)

Long Context benchmarks
BenchmarkInklingLlama 3.2 3B
LMArena Longer Query14341100

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Llama 3.2 3B: 24.7 (#307)

Writing & Preference benchmarks
BenchmarkInklingLlama 3.2 3B
LMArena Text14411110
LMArena Creative Writing13871094
EQ-Bench Creative Writing1611595
LMArena Multi-Turn14361105
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Llama 3.2 3B?

Inkling is the stronger model overall, scoring 44.1 to 28.9 on the Noometry Index. Llama 3.2 3B costs 21× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Which is cheaper, Inkling or Llama 3.2 3B?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Inkling lists at $1.87 and $4.68.

Is Inkling or Llama 3.2 3B better for coding?

Inkling scores higher on coding benchmarks: 34.5 versus 27.6 in the Noometry coding category.

Which has the bigger context window?

Llama 3.2 3B does, with 131K tokens against 66K.

How many benchmarks do Inkling and Llama 3.2 3B share?

14 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Llama 3.2 3B has 18.

Related comparisons

Go deeper