Model comparison

Claude Opus 4.1 vs Inkling-Small

Inkling-Small is the stronger model overall, scoring 46.5 to 41.0 on the Noometry Index.

Last verified . 25 shared benchmarks.

Claude Opus 4.1 Anthropic

41.0

Rank #142 Confirmed

Inkling-Small Thinking Machines Lab

46.5

Rank #63 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude Opus 4.1 scores higher in 5 categories and Inkling-Small in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Inkling-Small leads 45.1 to 22.3.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 12.6% for Claude Opus 4.1 and 46.3% for Inkling-Small.
  • Inkling-Small is cheaper at $0.45 / $1.20 per million input/output tokens, against $15 / $75 for Claude Opus 4.1.
  • Inkling-Small accepts more context: 524K tokens versus 200K.
  • Inkling-Small has downloadable open weights; the other is API-only.

Side by side

Claude Opus 4.1 and Inkling-Small specifications
Claude Opus 4.1Inkling-Small
ProviderAnthropicThinking Machines Lab
Noometry Index41.046.5
Released2025-08-052026-07-15
WeightsProprietaryOpen
Context window200K524K
Max output32K1.05M
Input $ / M tokens$15$0.45
Output $ / M tokens$75$1.20
Results tracked4833

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude Opus 4.1: 44.4 (#73), Inkling-Small: 43.6 (#85)

Coding benchmarks
BenchmarkClaude Opus 4.1Inkling-Small
LMArena WebDev13901409
LMArena Coding14791451
SWE-bench Verified73.3%—
SciCode—48.7%
WeirdML45.9%—
ALE-Bench674.77—
AlgoTune1.34—

Agentic & Tool Use Not comparable

Claude Opus 4.1: 35.0 (#41), Inkling-Small: —

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.1Inkling-Small
Terminal-Bench38%—
GDPval43.6%—
Cybench42%—
DeepResearch Bench48.3%—
LMArena Search1148—
METR Time Horizons66.8%—

Reasoning Inkling-Small leads

Claude Opus 4.1: 32.2 (#76), Inkling-Small: 38.6 (#63)

Reasoning benchmarks
BenchmarkClaude Opus 4.1Inkling-Small
Chess Puzzles7%18%
LMArena Hard Prompts14431423
Mystery Game Puzzles21%6%
Epoch Capabilities Index144.12150.15
ARC-AGI-2—40.1%
SimpleBench60%—
ARC-AGI-1—84%
CritPt—8.3%
EnigmaEval7.2%—
EBR-Bench7.9%—
DTBench80%—
LMCA37.1%—
ForecastBench62—

Math Inkling-Small leads

Claude Opus 4.1: 22.3 (#277), Inkling-Small: 45.1 (#77)

Math benchmarks
BenchmarkClaude Opus 4.1Inkling-Small
FrontierMath (Tiers 1-3)12.6%46.3%
FrontierMath Tier 42.4%17.1%
OTIS Mock AIME 2024-202568.9%90%
LMArena Math14311459
ProofBench—6%
FrontierMath (Feb 2025 set)7.2%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Inkling-Small leads

Claude Opus 4.1: 42.0 (#101), Inkling-Small: 48.2 (#77)

Knowledge benchmarks
BenchmarkClaude Opus 4.1Inkling-Small
GPQA Diamond77.3%88.5%
LMArena Expert14391442
Humanity's Last Exam11.5%—
SimpleQA Verified—19.1%
Confabulations17.1%—
Vectara Hallucination Rate11.8%—

Multimodal Inkling-Small leads

Claude Opus 4.1: 26.8 (#119), Inkling-Small: 39.1 (#62)

Multimodal benchmarks
BenchmarkClaude Opus 4.1Inkling-Small
LMArena Vision—1235
VPCT35%—

Multilingual Too close to call

Claude Opus 4.1: 52.0 (#95), Inkling-Small: 51.7 (#104)

Multilingual benchmarks
BenchmarkClaude Opus 4.1Inkling-Small
LMArena Non-English14051402
LMArena Chinese14271465
LMArena French14311436
LMArena German14131405
LMArena Japanese13781405
LMArena Korean13801363
LMArena Russian14221391
LMArena Spanish14481428

Instruction Following Claude Opus 4.1 leads

Claude Opus 4.1: 75.6 (#58), Inkling-Small: 73.8 (#114)

Instruction Following benchmarks
BenchmarkClaude Opus 4.1Inkling-Small
LMArena Instruction Following14351399

Long Context Claude Opus 4.1 leads

Claude Opus 4.1: 44.5 (#63), Inkling-Small: 42.7 (#118)

Long Context benchmarks
BenchmarkClaude Opus 4.1Inkling-Small
LMArena Longer Query14551401

Writing & Preference Claude Opus 4.1 leads

Claude Opus 4.1: 62.4 (#74), Inkling-Small: 59.6 (#107)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.1Inkling-Small
LMArena Text14191414
LMArena Creative Writing14121331
LMArena Multi-Turn14441418
Short-Story Creative Writing84.7%—
EQ-Bench Creative Writing—1491

Frequently asked questions

Is Claude Opus 4.1 better than Inkling-Small?

Inkling-Small is the stronger model overall, scoring 46.5 to 41.0 on the Noometry Index.

Which is cheaper, Claude Opus 4.1 or Inkling-Small?

Inkling-Small is cheaper. It lists at $0.45 per million input tokens and $1.20 per million output tokens; Claude Opus 4.1 lists at $15 and $75.

Is Claude Opus 4.1 or Inkling-Small better for coding?

They score almost the same on coding (44.4 vs 43.6); test both on your own repository before choosing.

Which has the bigger context window?

Inkling-Small does, with 524K tokens against 200K.

How many benchmarks do Claude Opus 4.1 and Inkling-Small share?

25 benchmarks have published results for both models. Claude Opus 4.1 has 48 scored results on Noometry and Inkling-Small has 33.

Related comparisons

Go deeper