Model comparison

Inkling vs Olmo 3.1 32b Instruct

Inkling is the stronger model overall, scoring 44.1 to 39.4 on the Noometry Index.

Last verified . 16 shared benchmarks.

Summary

  • They share 16 benchmarks with published results for both. Inkling scores higher in 6 categories and Olmo 3.1 32b Instruct in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 36.1.

Side by side

Inkling and Olmo 3.1 32b Instruct specifications
InklingOlmo 3.1 32b Instruct
ProviderThinking Machines LabAllen Institute for AI (Ai2)
Noometry Index44.139.4
Released2026-07-15—
WeightsOpenOpen
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked4116

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Inkling: 34.5 (#234), Olmo 3.1 32b Instruct: 39.5 (#157)

Coding benchmarks
BenchmarkInklingOlmo 3.1 32b Instruct
LMArena Coding14641347
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Olmo 3.1 32b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkInklingOlmo 3.1 32b Instruct
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Olmo 3.1 32b Instruct: 26.4 (#132)

Reasoning benchmarks
BenchmarkInklingOlmo 3.1 32b Instruct
LMArena Hard Prompts14511322
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Olmo 3.1 32b Instruct leads

Inkling: 31.3 (#225), Olmo 3.1 32b Instruct: 36.3 (#167)

Math benchmarks
BenchmarkInklingOlmo 3.1 32b Instruct
LMArena Math14791305
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Olmo 3.1 32b Instruct: 36.1 (#175)

Knowledge benchmarks
BenchmarkInklingOlmo 3.1 32b Instruct
LMArena Expert14651308
GPQA Diamond88.3%—
SimpleQA Verified40.3%—

Multilingual Inkling leads

Inkling: 54.0 (#52), Olmo 3.1 32b Instruct: 42.6 (#191)

Multilingual benchmarks
BenchmarkInklingOlmo 3.1 32b Instruct
LMArena Non-English14341275
LMArena Chinese14901304
LMArena French14581328
LMArena German14461282
LMArena Korean14041206
LMArena Russian14291268
LMArena Spanish14481336
LMArena Japanese1429—

Instruction Following Inkling leads

Inkling: 75.1 (#71), Olmo 3.1 32b Instruct: 68.6 (#187)

Instruction Following benchmarks
BenchmarkInklingOlmo 3.1 32b Instruct
LMArena Instruction Following14261299

Long Context Inkling leads

Inkling: 43.8 (#86), Olmo 3.1 32b Instruct: 39.9 (#166)

Long Context benchmarks
BenchmarkInklingOlmo 3.1 32b Instruct
LMArena Longer Query14341312

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Olmo 3.1 32b Instruct: 50.2 (#185)

Writing & Preference benchmarks
BenchmarkInklingOlmo 3.1 32b Instruct
LMArena Text14411311
LMArena Creative Writing13871264
LMArena Multi-Turn14361309
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Olmo 3.1 32b Instruct?

Inkling is the stronger model overall, scoring 44.1 to 39.4 on the Noometry Index.

Is Inkling or Olmo 3.1 32b Instruct better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 34.5 in the Noometry coding category.

How many benchmarks do Inkling and Olmo 3.1 32b Instruct share?

16 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Olmo 3.1 32b Instruct has 16.

Related comparisons

Go deeper