Model comparison

Inkling vs Mercury

Inkling is the stronger model overall, scoring 44.1 to 37.6 on the Noometry Index.

Last verified . 8 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Inkling scores higher in 5 categories and Mercury in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Inkling leads 40.4 to 17.5.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Inkling and Mercury specifications
InklingMercury
ProviderThinking Machines LabInception
Noometry Index44.137.6
Released2026-07-15—
WeightsOpenProprietary
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked419

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Inkling: 34.5 (#234), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkInklingMercury
LMArena Coding14641322
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Mercury: —

Agentic & Tool Use benchmarks
BenchmarkInklingMercury
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkInklingMercury
LMArena Hard Prompts14511285
ARC-AGI-236.5%—
SimpleBench50%—
Kagi LLM Benchmark—21.6%
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Not comparable

Inkling: 31.3 (#225), Mercury: —

Math benchmarks
BenchmarkInklingMercury
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—
LMArena Math1479—

Knowledge Not comparable

Inkling: 55.1 (#49), Mercury: —

Knowledge benchmarks
BenchmarkInklingMercury
GPQA Diamond88.3%—
SimpleQA Verified40.3%—
LMArena Expert1465—

Multilingual Inkling leads

Inkling: 54.0 (#52), Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkInklingMercury
LMArena Non-English14341260
LMArena Chinese1490—
LMArena French1458—
LMArena German1446—
LMArena Japanese1429—
LMArena Korean1404—
LMArena Russian1429—
LMArena Spanish1448—

Instruction Following Inkling leads

Inkling: 75.1 (#71), Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkInklingMercury
LMArena Instruction Following14261239

Long Context Inkling leads

Inkling: 43.8 (#86), Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkInklingMercury
LMArena Longer Query14341266

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkInklingMercury
LMArena Text14411282
LMArena Creative Writing13871191
LMArena Multi-Turn14361282
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Mercury?

Inkling is the stronger model overall, scoring 44.1 to 37.6 on the Noometry Index.

Is Inkling or Mercury better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 34.5 in the Noometry coding category.

How many benchmarks do Inkling and Mercury share?

8 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper