Model comparison

Inkling vs Olmo 3.1 32b Think

Inkling is the stronger model overall, scoring 44.1 to 37.9 on the Noometry Index.

Last verified . 15 shared benchmarks.

Summary

  • They share 15 benchmarks with published results for both. Inkling scores higher in 6 categories and Olmo 3.1 32b Think in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 35.7.

Side by side

Inkling and Olmo 3.1 32b Think specifications
InklingOlmo 3.1 32b Think
ProviderThinking Machines LabAllen Institute for AI (Ai2)
Noometry Index44.137.9
Released2026-07-15—
WeightsOpenOpen
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked4115

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Think leads

Inkling: 34.5 (#234), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkInklingOlmo 3.1 32b Think
LMArena Coding14641291
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Olmo 3.1 32b Think: —

Agentic & Tool Use benchmarks
BenchmarkInklingOlmo 3.1 32b Think
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkInklingOlmo 3.1 32b Think
LMArena Hard Prompts14511272
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Olmo 3.1 32b Think leads

Inkling: 31.3 (#225), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkInklingOlmo 3.1 32b Think
LMArena Math14791305
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkInklingOlmo 3.1 32b Think
LMArena Expert14651295
GPQA Diamond88.3%—
SimpleQA Verified40.3%—

Multilingual Inkling leads

Inkling: 54.0 (#52), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkInklingOlmo 3.1 32b Think
LMArena Non-English14341209
LMArena Chinese14901242
LMArena French14581260
LMArena German14461262
LMArena Russian14291193
LMArena Spanish14481289
LMArena Japanese1429—
LMArena Korean1404—

Instruction Following Inkling leads

Inkling: 75.1 (#71), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkInklingOlmo 3.1 32b Think
LMArena Instruction Following14261247

Long Context Inkling leads

Inkling: 43.8 (#86), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkInklingOlmo 3.1 32b Think
LMArena Longer Query14341272

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkInklingOlmo 3.1 32b Think
LMArena Text14411272
LMArena Creative Writing13871226
LMArena Multi-Turn14361252
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Olmo 3.1 32b Think?

Inkling is the stronger model overall, scoring 44.1 to 37.9 on the Noometry Index.

Is Inkling or Olmo 3.1 32b Think better for coding?

Olmo 3.1 32b Think scores higher on coding benchmarks: 37.7 versus 34.5 in the Noometry coding category.

How many benchmarks do Inkling and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper