Model comparison

Inkling vs Qwen Max

Inkling is the stronger model overall, scoring 44.1 to 34.7 on the Noometry Index.

Last verified . 19 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Inkling scores higher in 8 categories and Qwen Max in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 30.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 88.9% for Inkling and 16.1% for Qwen Max.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Inkling accepts more context: 66K tokens versus 33K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Inkling and Qwen Max specifications
InklingQwen Max
ProviderThinking Machines LabAlibaba (Qwen)
Noometry Index44.134.7
Released2026-07-152024-04-03
WeightsOpenProprietary
Context window66K33K
Max output66K8K
Input $ / M tokens$1.87$1.60
Output $ / M tokens$4.68$6.40
Results tracked4123

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling leads

Inkling: 34.5 (#234), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkInklingQwen Max
LMArena Coding14641288
FrontierCode14%—
Aider Polyglot—21.8%
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkInklingQwen Max
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkInklingQwen Max
LMArena Hard Prompts14511269
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Inkling leads

Inkling: 31.3 (#225), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkInklingQwen Max
OTIS Mock AIME 2024-202588.9%16.1%
LMArena Math14791275
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
ProofBench0%—
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Inkling leads

Inkling: 55.1 (#49), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkInklingQwen Max
GPQA Diamond88.3%56.1%
LMArena Expert14651248
SimpleQA Verified40.3%—

Multilingual Inkling leads

Inkling: 54.0 (#52), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkInklingQwen Max
LMArena Non-English14341263
LMArena Chinese14901254
LMArena French14581330
LMArena German14461254
LMArena Japanese14291205
LMArena Korean14041142
LMArena Russian14291274
LMArena Spanish14481290

Instruction Following Inkling leads

Inkling: 75.1 (#71), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkInklingQwen Max
LMArena Instruction Following14261262

Long Context Inkling leads

Inkling: 43.8 (#86), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkInklingQwen Max
LMArena Longer Query14341288
Fiction.LiveBench—66.7%

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkInklingQwen Max
LMArena Text14411282
LMArena Creative Writing13871248
LMArena Multi-Turn14361277
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Qwen Max?

Inkling is the stronger model overall, scoring 44.1 to 34.7 on the Noometry Index.

Which is cheaper, Inkling or Qwen Max?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Inkling or Qwen Max better for coding?

Inkling scores higher on coding benchmarks: 34.5 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Inkling does, with 66K tokens against 33K.

How many benchmarks do Inkling and Qwen Max share?

19 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper