Model comparison

Inkling vs Qwen3.6 Max Preview

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 44.1 on the Noometry Index.

Last verified . 23 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Inkling scores higher in 1 category and Qwen3.6 Max Preview in 7 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.6 Max Preview leads 54.1 to 31.3.
  • The biggest single-benchmark swing is SimpleBench: 50% for Inkling and 63% for Qwen3.6 Max Preview.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $1.30 / $7.80 for Qwen3.6 Max Preview.
  • Qwen3.6 Max Preview accepts more context: 262K tokens versus 66K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Inkling and Qwen3.6 Max Preview specifications
InklingQwen3.6 Max Preview
ProviderThinking Machines LabAlibaba (Qwen)
Noometry Index44.151.5
Released2026-07-152026-04-20
WeightsOpenProprietary
Context window66K262K
Max output66K66K
Input $ / M tokens$1.87$1.30
Output $ / M tokens$4.68$7.80
Results tracked4129

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

Inkling: 34.5 (#234), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkInklingQwen3.6 Max Preview
LMArena WebDev14131482
LMArena Coding14641471
SWE-bench Verified—76.7%
FrontierCode14%—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkInklingQwen3.6 Max Preview
APEX-Agents33.8%—
τ²-bench Banking25%—
Vending-Bench 2—4,254

Reasoning Qwen3.6 Max Preview leads

Inkling: 40.4 (#56), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkInklingQwen3.6 Max Preview
SimpleBench50%63%
Chess Puzzles21%20%
LMArena Hard Prompts14511457
DTBench87.5%87.2%
LMCA37.6%42.5%
Epoch Capabilities Index148.54149.24
ARC-AGI-236.5%—
NYT Connections (extended)—74.1%
ARC-AGI-179.5%—
CritPt5.4%—
Mystery Game Puzzles—19%

Math Qwen3.6 Max Preview leads

Inkling: 31.3 (#225), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkInklingQwen3.6 Max Preview
OTIS Mock AIME 2024-202588.9%91.1%
LMArena Math14791465
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
ProofBench0%—
FrontierMath (Feb 2025 set)—23.1%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Qwen3.6 Max Preview leads

Inkling: 55.1 (#49), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkInklingQwen3.6 Max Preview
GPQA Diamond88.3%87.4%
SimpleQA Verified40.3%52%
LMArena Expert14651478

Multilingual Too close to call

Inkling: 54.0 (#52), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkInklingQwen3.6 Max Preview
LMArena Non-English14341437
LMArena Chinese14901487
LMArena French14581449
LMArena Russian14291445
LMArena Spanish14481454
LMArena German1446—
LMArena Japanese1429—
LMArena Korean1404—

Instruction Following Too close to call

Inkling: 75.1 (#71), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkInklingQwen3.6 Max Preview
LMArena Instruction Following14261438

Long Context Too close to call

Inkling: 43.8 (#86), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkInklingQwen3.6 Max Preview
LMArena Longer Query14341457

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkInklingQwen3.6 Max Preview
LMArena Text14411447
LMArena Creative Writing13871435
LMArena Multi-Turn14361456
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Qwen3.6 Max Preview?

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 44.1 on the Noometry Index.

Which is cheaper, Inkling or Qwen3.6 Max Preview?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; Qwen3.6 Max Preview lists at $1.30 and $7.80.

Is Inkling or Qwen3.6 Max Preview better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

Qwen3.6 Max Preview does, with 262K tokens against 66K.

How many benchmarks do Inkling and Qwen3.6 Max Preview share?

23 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper