Model comparison

Claude 3.5 Haiku vs Inkling

Inkling is the stronger model overall, scoring 44.1 to 29.2 on the Noometry Index.

Last verified . 25 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and Inkling in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 18.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.3% for Claude 3.5 Haiku and 88.9% for Inkling.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Inkling specifications
Claude 3.5 HaikuInkling
ProviderAnthropicThinking Machines Lab
Noometry Index29.244.1
Released2024-10-222026-07-15
WeightsProprietaryOpen
Context window—66K
Max output—66K
Input $ / M tokens—$1.87
Output $ / M tokens—$4.68
Results tracked4941

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling leads

Claude 3.5 Haiku: 32.9 (#265), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkClaude 3.5 HaikuInkling
SciCode27.4%47%
WeirdML30.7%32.3%
LMArena Coding12861464
FrontierCode—14%
Aider Polyglot28%—
LMArena WebDev—1413
FrontierSWE—4.1%
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—
ALE-Bench—946

Agentic & Tool Use Inkling leads

Claude 3.5 Haiku: 28.0 (#95), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuInkling
APEX-Agents—33.8%
τ²-bench Banking—25%
BALROG19.3%—

Reasoning Inkling leads

Claude 3.5 Haiku: 17.7 (#290), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuInkling
CritPt0%5.4%
LMArena Hard Prompts12511451
DTBench56.7%87.5%
Epoch Capabilities Index127.15148.54
ARC-AGI-2—36.5%
SimpleBench—50%
ARC-AGI-1—79.5%
Chess Puzzles—21%
LiveBench Reasoning28.1%—
LiveBench Data Analysis48.5%—
LMCA—37.6%
LiveBench43.5%—

Math Inkling leads

Claude 3.5 Haiku: 14.7 (#300), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkClaude 3.5 HaikuInkling
OTIS Mock AIME 2024-20254.3%88.9%
LMArena Math12441479
FrontierMath (Tiers 1-3)—33.3%
FrontierMath Tier 4—4.9%
ProofBench—0%
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Inkling leads

Claude 3.5 Haiku: 18.7 (#281), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuInkling
GPQA Diamond38.1%88.3%
LMArena Expert12081465
SimpleQA Verified—40.3%
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Inkling: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuInkling
LMArena Vision1092—
GeoBench34%—

Multilingual Inkling leads

Claude 3.5 Haiku: 40.0 (#218), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuInkling
LMArena Non-English12381434
LMArena Chinese12291490
LMArena French12641458
LMArena German12371446
LMArena Japanese11751429
LMArena Korean11731404
LMArena Russian12531429
LMArena Spanish12611448

Instruction Following Inkling leads

Claude 3.5 Haiku: 62.9 (#234), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuInkling
LMArena Instruction Following12411426
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Inkling leads

Claude 3.5 Haiku: 38.3 (#200), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuInkling
LMArena Longer Query12611434

Writing & Preference Inkling leads

Claude 3.5 Haiku: 42.7 (#234), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuInkling
LMArena Text12551441
LMArena Creative Writing12331387
EQ-Bench Creative Writing11461611
LMArena Multi-Turn12651436
Short-Story Creative Writing73.5%—
WildBench76%—
EQ-Bench 4—1226
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Inkling?

Inkling is the stronger model overall, scoring 44.1 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Inkling better for coding?

Inkling scores higher on coding benchmarks: 34.5 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Inkling share?

25 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper