Model comparison

Claude 3.5 Sonnet vs Inkling

Inkling is the stronger model overall, scoring 44.1 to 34.6 on the Noometry Index.

Last verified . 24 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 2 categories and Inkling in 7 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 28.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.5% for Claude 3.5 Sonnet and 88.9% for Inkling.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Inkling specifications
Claude 3.5 SonnetInkling
ProviderAnthropicThinking Machines Lab
Noometry Index34.644.1
Released2024-06-202026-07-15
WeightsProprietaryOpen
Context window—66K
Max output—66K
Input $ / M tokens—$1.87
Output $ / M tokens—$4.68
Results tracked6041

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkClaude 3.5 SonnetInkling
WeirdML40%32.3%
LMArena Coding13421464
FrontierCode—14%
Aider Polyglot51.6%—
LMArena WebDev—1413
FrontierSWE—4.1%
SciCode—47%
GSO4.6%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
ALE-Bench—946
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 32.3 (#67), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetInkling
APEX-Agents—33.8%
TheAgentCompany24%—
τ²-bench Banking—25%
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Inkling leads

Claude 3.5 Sonnet: 23.1 (#183), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetInkling
SimpleBench41.4%50%
LMArena Hard Prompts13051451
DTBench67.8%87.5%
Epoch Capabilities Index133.55148.54
ARC-AGI-2—36.5%
ARC-AGI-1—79.5%
CritPt—5.4%
Chess Puzzles—21%
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
LiveBench Data Analysis55%—
LMCA—37.6%
ForecastBench60.7—
LiveBench59%—

Math Inkling leads

Claude 3.5 Sonnet: 19.2 (#288), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkClaude 3.5 SonnetInkling
OTIS Mock AIME 2024-20258.5%88.9%
LMArena Math13071479
FrontierMath (Tiers 1-3)—33.3%
FrontierMath Tier 4—4.9%
ProofBench—0%
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Inkling leads

Claude 3.5 Sonnet: 28.6 (#245), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetInkling
GPQA Diamond55.3%88.3%
LMArena Expert12651465
Humanity's Last Exam4.1%—
SimpleQA Verified—40.3%
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Inkling: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetInkling
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Inkling leads

Claude 3.5 Sonnet: 43.2 (#185), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetInkling
LMArena Non-English12831434
LMArena Chinese12721490
LMArena French13051458
LMArena German12971446
LMArena Japanese12341429
LMArena Korean12001404
LMArena Russian13061429
LMArena Spanish12901448

Instruction Following Inkling leads

Claude 3.5 Sonnet: 68.8 (#182), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetInkling
LMArena Instruction Following12971426
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Inkling leads

Claude 3.5 Sonnet: 39.9 (#167), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetInkling
LMArena Longer Query13111434

Writing & Preference Inkling leads

Claude 3.5 Sonnet: 52.9 (#164), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetInkling
LMArena Text12981441
LMArena Creative Writing12921387
EQ-Bench Creative Writing14511611
LMArena Multi-Turn13261436
Short-Story Creative Writing80.3%—
WildBench79.2%—
EQ-Bench 4—1226
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Inkling?

Inkling is the stronger model overall, scoring 44.1 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Inkling better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 34.5 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Inkling share?

24 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper