Model comparison

Inkling vs Qwen2.5 Plus 1127

Inkling is the stronger model overall, scoring 44.1 to 38.8 on the Noometry Index.

Last verified . 14 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Inkling scores higher in 6 categories and Qwen2.5 Plus 1127 in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 35.5.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Inkling and Qwen2.5 Plus 1127 specifications
InklingQwen2.5 Plus 1127
ProviderThinking Machines LabAlibaba (Qwen)
Noometry Index44.138.8
Released2026-07-15—
WeightsOpenProprietary
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked4114

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Inkling: 34.5 (#234), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkInklingQwen2.5 Plus 1127
LMArena Coding14641314
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Qwen2.5 Plus 1127: —

Agentic & Tool Use benchmarks
BenchmarkInklingQwen2.5 Plus 1127
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkInklingQwen2.5 Plus 1127
LMArena Hard Prompts14511299
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Qwen2.5 Plus 1127 leads

Inkling: 31.3 (#225), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkInklingQwen2.5 Plus 1127
LMArena Math14791298
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkInklingQwen2.5 Plus 1127
LMArena Expert14651289
GPQA Diamond88.3%—
SimpleQA Verified40.3%—

Multilingual Inkling leads

Inkling: 54.0 (#52), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkInklingQwen2.5 Plus 1127
LMArena Non-English14341265
LMArena Chinese14901314
LMArena German14461231
LMArena Japanese14291207
LMArena Russian14291271
LMArena French1458—
LMArena Korean1404—
LMArena Spanish1448—

Instruction Following Inkling leads

Inkling: 75.1 (#71), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkInklingQwen2.5 Plus 1127
LMArena Instruction Following14261275

Long Context Inkling leads

Inkling: 43.8 (#86), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkInklingQwen2.5 Plus 1127
LMArena Longer Query14341292

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkInklingQwen2.5 Plus 1127
LMArena Text14411299
LMArena Creative Writing13871262
LMArena Multi-Turn14361299
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Qwen2.5 Plus 1127?

Inkling is the stronger model overall, scoring 44.1 to 38.8 on the Noometry Index.

Is Inkling or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 34.5 in the Noometry coding category.

How many benchmarks do Inkling and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper