Model comparison

Inkling vs Qwen3.5-Flash

Inkling is the stronger model overall, scoring 44.1 to 42.5 on the Noometry Index. Qwen3.5-Flash costs 15× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Last verified . 27 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Qwen3.5-Flash Alibaba (Qwen)

42.5

Rank #112 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Inkling scores higher in 7 categories and Qwen3.5-Flash in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 43.2.
  • The biggest single-benchmark swing is SimpleQA Verified: 40.3% for Inkling and 20.3% for Qwen3.5-Flash.
  • Qwen3.5-Flash is cheaper at $0.10 / $0.40 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • Qwen3.5-Flash accepts more context: 1M tokens versus 66K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Inkling and Qwen3.5-Flash specifications
InklingQwen3.5-Flash
ProviderThinking Machines LabAlibaba (Qwen)
Noometry Index44.142.5
Released2026-07-152026-02-23
WeightsOpenProprietary
Context window66K1M
Max output66K66K
Input $ / M tokens$1.87$0.10
Output $ / M tokens$4.68$0.40
Results tracked4132

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Inkling: 34.5 (#234), Qwen3.5-Flash: 34.2 (#242)

Coding benchmarks
BenchmarkInklingQwen3.5-Flash
LMArena WebDev14131244
LMArena Coding14641412
ALE-Bench946221.8
FrontierCode14%—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Qwen3.5-Flash: —

Agentic & Tool Use benchmarks
BenchmarkInklingQwen3.5-Flash
APEX-Agents33.8%—
τ²-bench Banking25%—
Vending-Bench 2—462.69

Reasoning Inkling leads

Inkling: 40.4 (#56), Qwen3.5-Flash: 33.7 (#72)

Reasoning benchmarks
BenchmarkInklingQwen3.5-Flash
Chess Puzzles21%21%
LMArena Hard Prompts14511403
DTBench87.5%82.9%
LMCA37.6%29.1%
Epoch Capabilities Index148.54143.98
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Mystery Game Puzzles—20%

Math Qwen3.5-Flash leads

Inkling: 31.3 (#225), Qwen3.5-Flash: 37.4 (#158)

Math benchmarks
BenchmarkInklingQwen3.5-Flash
FrontierMath (Tiers 1-3)33.3%18.2%
OTIS Mock AIME 2024-202588.9%84.4%
LMArena Math14791407
FrontierMath Tier 44.9%—
ProofBench0%—
FrontierMath (Feb 2025 set)—6.2%
FrontierMath Tier 4 (v1)—0%

Knowledge Inkling leads

Inkling: 55.1 (#49), Qwen3.5-Flash: 43.2 (#93)

Knowledge benchmarks
BenchmarkInklingQwen3.5-Flash
GPQA Diamond88.3%82.3%
SimpleQA Verified40.3%20.3%
LMArena Expert14651407
Vectara Hallucination Rate—10.5%

Multilingual Inkling leads

Inkling: 54.0 (#52), Qwen3.5-Flash: 50.5 (#121)

Multilingual benchmarks
BenchmarkInklingQwen3.5-Flash
LMArena Non-English14341385
LMArena Chinese14901446
LMArena French14581412
LMArena German14461390
LMArena Japanese14291368
LMArena Korean14041344
LMArena Russian14291379
LMArena Spanish14481400

Instruction Following Inkling leads

Inkling: 75.1 (#71), Qwen3.5-Flash: 72.6 (#139)

Instruction Following benchmarks
BenchmarkInklingQwen3.5-Flash
LMArena Instruction Following14261374

Long Context Inkling leads

Inkling: 43.8 (#86), Qwen3.5-Flash: 42.4 (#124)

Long Context benchmarks
BenchmarkInklingQwen3.5-Flash
LMArena Longer Query14341392

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Qwen3.5-Flash: 57.9 (#122)

Writing & Preference benchmarks
BenchmarkInklingQwen3.5-Flash
LMArena Text14411397
LMArena Creative Writing13871343
LMArena Multi-Turn14361393
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Qwen3.5-Flash?

Inkling is the stronger model overall, scoring 44.1 to 42.5 on the Noometry Index. Qwen3.5-Flash costs 15× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Which is cheaper, Inkling or Qwen3.5-Flash?

Qwen3.5-Flash is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; Inkling lists at $1.87 and $4.68.

Is Inkling or Qwen3.5-Flash better for coding?

They score almost the same on coding (34.5 vs 34.2); test both on your own repository before choosing.

Which has the bigger context window?

Qwen3.5-Flash does, with 1M tokens against 66K.

How many benchmarks do Inkling and Qwen3.5-Flash share?

27 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Qwen3.5-Flash has 32.

Related comparisons

Go deeper