Model comparison

Inkling vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 44.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Inkling scores higher in 2 categories and Qwen3.5 Max Preview in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 41.8.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Inkling and Qwen3.5 Max Preview specifications
InklingQwen3.5 Max Preview
ProviderThinking Machines LabAlibaba (Qwen)
Noometry Index44.145.3
Released2026-07-15—
WeightsOpenProprietary
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked4117

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Inkling: 34.5 (#234), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkInklingQwen3.5 Max Preview
LMArena Coding14641487
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Qwen3.5 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkInklingQwen3.5 Max Preview
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkInklingQwen3.5 Max Preview
LMArena Hard Prompts14511483
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Qwen3.5 Max Preview leads

Inkling: 31.3 (#225), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkInklingQwen3.5 Max Preview
LMArena Math14791474
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkInklingQwen3.5 Max Preview
LMArena Expert14651489
GPQA Diamond88.3%—
SimpleQA Verified40.3%—

Multilingual Qwen3.5 Max Preview leads

Inkling: 54.0 (#52), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkInklingQwen3.5 Max Preview
LMArena Non-English14341465
LMArena Chinese14901534
LMArena French14581484
LMArena German14461487
LMArena Japanese14291495
LMArena Korean14041438
LMArena Russian14291471
LMArena Spanish14481470

Instruction Following Qwen3.5 Max Preview leads

Inkling: 75.1 (#71), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkInklingQwen3.5 Max Preview
LMArena Instruction Following14261467

Long Context Qwen3.5 Max Preview leads

Inkling: 43.8 (#86), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkInklingQwen3.5 Max Preview
LMArena Longer Query14341476

Writing & Preference Too close to call

Inkling: 65.2 (#51), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkInklingQwen3.5 Max Preview
LMArena Text14411470
LMArena Creative Writing13871464
LMArena Multi-Turn14361478
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 44.1 on the Noometry Index.

Is Inkling or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 34.5 in the Noometry coding category.

How many benchmarks do Inkling and Qwen3.5 Max Preview share?

17 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper