Model comparison

Inkling vs Llama 3.2 90B

Inkling is the stronger model overall, scoring 44.1 to 27.5 on the Noometry Index.

Last verified . 3 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Inkling scores higher in 3 categories and Llama 3.2 90B in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 21.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 88.9% for Inkling and 2.6% for Llama 3.2 90B.

Side by side

Inkling and Llama 3.2 90B specifications
InklingLlama 3.2 90B
ProviderThinking Machines LabMeta
Noometry Index44.127.5
Released2026-07-152024-09-24
WeightsOpenOpen
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked419

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Inkling: 34.5 (#234), Llama 3.2 90B: —

Coding benchmarks
BenchmarkInklingLlama 3.2 90B
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
LMArena Coding1464—
ALE-Bench946—

Agentic & Tool Use Too close to call

Inkling: 29.6 (#85), Llama 3.2 90B: 30.0 (#80)

Agentic & Tool Use benchmarks
BenchmarkInklingLlama 3.2 90B
APEX-Agents33.8%—
τ²-bench Banking25%—
BALROG—27.3%

Reasoning Inkling leads

Inkling: 40.4 (#56), Llama 3.2 90B: 21.7 (#217)

Reasoning benchmarks
BenchmarkInklingLlama 3.2 90B
Epoch Capabilities Index148.54125.5
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
EnigmaEval—0.4%
LMArena Hard Prompts1451—
DTBench87.5%—
LMCA37.6%—

Math Inkling leads

Inkling: 31.3 (#225), Llama 3.2 90B: 11.1 (#308)

Math benchmarks
BenchmarkInklingLlama 3.2 90B
OTIS Mock AIME 2024-202588.9%2.6%
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
ProofBench0%—
LMArena Math1479—
MATH Level 5—39.4%

Knowledge Inkling leads

Inkling: 55.1 (#49), Llama 3.2 90B: 21.7 (#274)

Knowledge benchmarks
BenchmarkInklingLlama 3.2 90B
GPQA Diamond88.3%41%
SimpleQA Verified40.3%—
LMArena Expert1465—
MMLU—80.3%

Multimodal Not comparable

Inkling: —, Llama 3.2 90B: 25.4 (#124)

Multimodal benchmarks
BenchmarkInklingLlama 3.2 90B
LMArena Vision—1000
GeoBench—52%

Multilingual Not comparable

Inkling: 54.0 (#52), Llama 3.2 90B: —

Multilingual benchmarks
BenchmarkInklingLlama 3.2 90B
LMArena Non-English1434—
LMArena Chinese1490—
LMArena French1458—
LMArena German1446—
LMArena Japanese1429—
LMArena Korean1404—
LMArena Russian1429—
LMArena Spanish1448—

Instruction Following Not comparable

Inkling: 75.1 (#71), Llama 3.2 90B: —

Instruction Following benchmarks
BenchmarkInklingLlama 3.2 90B
LMArena Instruction Following1426—

Long Context Not comparable

Inkling: 43.8 (#86), Llama 3.2 90B: —

Long Context benchmarks
BenchmarkInklingLlama 3.2 90B
LMArena Longer Query1434—

Writing & Preference Not comparable

Inkling: 65.2 (#51), Llama 3.2 90B: —

Writing & Preference benchmarks
BenchmarkInklingLlama 3.2 90B
LMArena Text1441—
LMArena Creative Writing1387—
EQ-Bench Creative Writing1611—
EQ-Bench 41226—
LMArena Multi-Turn1436—

Frequently asked questions

Is Inkling better than Llama 3.2 90B?

Inkling is the stronger model overall, scoring 44.1 to 27.5 on the Noometry Index.

How many benchmarks do Inkling and Llama 3.2 90B share?

3 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Llama 3.2 90B has 9.

Related comparisons

Go deeper