Model comparison

Claude 2.1 vs Inkling

Inkling is the stronger model overall, scoring 44.1 to 25.2 on the Noometry Index.

Last verified . 5 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and Inkling in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 15.4.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 1.9% for Claude 2.1 and 88.9% for Inkling.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and Inkling specifications
Claude 2.1Inkling
ProviderAnthropicThinking Machines Lab
Noometry Index25.244.1
Released2023-11-212026-07-15
WeightsProprietaryOpen
Context window—66K
Max output—66K
Input $ / M tokens—$1.87
Output $ / M tokens—$4.68
Results tracked741

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling leads

Claude 2.1: 26.2 (#327), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkClaude 2.1Inkling
WeirdML7.1%32.3%
FrontierCode—14%
LMArena WebDev—1413
FrontierSWE—4.1%
SciCode—47%
LMArena Coding—1464
ALE-Bench—946

Agentic & Tool Use Not comparable

Claude 2.1: —, Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkClaude 2.1Inkling
APEX-Agents—33.8%
τ²-bench Banking—25%

Reasoning Inkling leads

Claude 2.1: 21.4 (#221), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkClaude 2.1Inkling
DTBench51%87.5%
Epoch Capabilities Index119.27148.54
ARC-AGI-2—36.5%
SimpleBench—50%
ARC-AGI-1—79.5%
CritPt—5.4%
Chess Puzzles—21%
LMArena Hard Prompts—1451
LMCA—37.6%
ForecastBench54.2—

Math Inkling leads

Claude 2.1: 10.2 (#315), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkClaude 2.1Inkling
OTIS Mock AIME 2024-20251.9%88.9%
FrontierMath (Tiers 1-3)—33.3%
FrontierMath Tier 4—4.9%
ProofBench—0%
LMArena Math—1479

Knowledge Inkling leads

Claude 2.1: 15.4 (#292), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkClaude 2.1Inkling
GPQA Diamond33%88.3%
SimpleQA Verified—40.3%
LMArena Expert—1465
MMLU73.5%—

Multilingual Not comparable

Claude 2.1: —, Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkClaude 2.1Inkling
LMArena Non-English—1434
LMArena Chinese—1490
LMArena French—1458
LMArena German—1446
LMArena Japanese—1429
LMArena Korean—1404
LMArena Russian—1429
LMArena Spanish—1448

Instruction Following Not comparable

Claude 2.1: —, Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkClaude 2.1Inkling
LMArena Instruction Following—1426

Long Context Not comparable

Claude 2.1: —, Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkClaude 2.1Inkling
LMArena Longer Query—1434

Writing & Preference Not comparable

Claude 2.1: —, Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkClaude 2.1Inkling
LMArena Text—1441
LMArena Creative Writing—1387
EQ-Bench Creative Writing—1611
EQ-Bench 4—1226
LMArena Multi-Turn—1436

Frequently asked questions

Is Claude 2.1 better than Inkling?

Inkling is the stronger model overall, scoring 44.1 to 25.2 on the Noometry Index.

Is Claude 2.1 or Inkling better for coding?

Inkling scores higher on coding benchmarks: 34.5 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Inkling share?

5 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper