Model comparison

Inkling-Small vs Qwen3.5-Flash

Inkling-Small is the stronger model overall, scoring 46.5 to 42.5 on the Noometry Index. Qwen3.5-Flash costs 3.6× less per token, which makes it the better buy when Inkling-Small's lead doesn't matter for your workload.

Last verified . 25 shared benchmarks.

Inkling-Small Thinking Machines Lab

46.5

Rank #63 Confirmed

Qwen3.5-Flash Alibaba (Qwen)

42.5

Rank #112 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Inkling-Small scores higher in 8 categories and Qwen3.5-Flash in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Inkling-Small leads 43.6 to 34.2.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 46.3% for Inkling-Small and 18.2% for Qwen3.5-Flash.
  • Qwen3.5-Flash is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.45 / $1.20 for Inkling-Small.
  • Qwen3.5-Flash accepts more context: 1M tokens versus 524K.
  • Inkling-Small has downloadable open weights; the other is API-only.

Side by side

Inkling-Small and Qwen3.5-Flash specifications
Inkling-SmallQwen3.5-Flash
ProviderThinking Machines LabAlibaba (Qwen)
Noometry Index46.542.5
Released2026-07-152026-02-23
WeightsOpenProprietary
Context window524K1M
Max output1.05M66K
Input $ / M tokens$0.45$0.10
Output $ / M tokens$1.20$0.40
Results tracked3332

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling-Small leads

Inkling-Small: 43.6 (#85), Qwen3.5-Flash: 34.2 (#242)

Coding benchmarks
BenchmarkInkling-SmallQwen3.5-Flash
LMArena WebDev14091244
LMArena Coding14511412
SciCode48.7%—
ALE-Bench—221.8

Agentic & Tool Use Not comparable

Inkling-Small: —, Qwen3.5-Flash: —

Agentic & Tool Use benchmarks
BenchmarkInkling-SmallQwen3.5-Flash
Vending-Bench 2—462.69

Reasoning Inkling-Small leads

Inkling-Small: 38.6 (#63), Qwen3.5-Flash: 33.7 (#72)

Reasoning benchmarks
BenchmarkInkling-SmallQwen3.5-Flash
Chess Puzzles18%21%
LMArena Hard Prompts14231403
Mystery Game Puzzles6%20%
Epoch Capabilities Index150.15143.98
ARC-AGI-240.1%—
ARC-AGI-184%—
CritPt8.3%—
DTBench—82.9%
LMCA—29.1%

Math Inkling-Small leads

Inkling-Small: 45.1 (#77), Qwen3.5-Flash: 37.4 (#158)

Math benchmarks
BenchmarkInkling-SmallQwen3.5-Flash
FrontierMath (Tiers 1-3)46.3%18.2%
OTIS Mock AIME 2024-202590%84.4%
LMArena Math14591407
FrontierMath Tier 417.1%—
ProofBench6%—
FrontierMath (Feb 2025 set)—6.2%
FrontierMath Tier 4 (v1)—0%

Knowledge Inkling-Small leads

Inkling-Small: 48.2 (#77), Qwen3.5-Flash: 43.2 (#93)

Knowledge benchmarks
BenchmarkInkling-SmallQwen3.5-Flash
GPQA Diamond88.5%82.3%
SimpleQA Verified19.1%20.3%
LMArena Expert14421407
Vectara Hallucination Rate—10.5%

Multimodal Not comparable

Inkling-Small: 39.1 (#62), Qwen3.5-Flash: —

Multimodal benchmarks
BenchmarkInkling-SmallQwen3.5-Flash
LMArena Vision1235—

Multilingual Inkling-Small leads

Inkling-Small: 51.7 (#104), Qwen3.5-Flash: 50.5 (#121)

Multilingual benchmarks
BenchmarkInkling-SmallQwen3.5-Flash
LMArena Non-English14021385
LMArena Chinese14651446
LMArena French14361412
LMArena German14051390
LMArena Japanese14051368
LMArena Korean13631344
LMArena Russian13911379
LMArena Spanish14281400

Instruction Following Inkling-Small leads

Inkling-Small: 73.8 (#114), Qwen3.5-Flash: 72.6 (#139)

Instruction Following benchmarks
BenchmarkInkling-SmallQwen3.5-Flash
LMArena Instruction Following13991374

Long Context Too close to call

Inkling-Small: 42.7 (#118), Qwen3.5-Flash: 42.4 (#124)

Long Context benchmarks
BenchmarkInkling-SmallQwen3.5-Flash
LMArena Longer Query14011392

Writing & Preference Inkling-Small leads

Inkling-Small: 59.6 (#107), Qwen3.5-Flash: 57.9 (#122)

Writing & Preference benchmarks
BenchmarkInkling-SmallQwen3.5-Flash
LMArena Text14141397
LMArena Creative Writing13311343
LMArena Multi-Turn14181393
EQ-Bench Creative Writing1491—

Frequently asked questions

Is Inkling-Small better than Qwen3.5-Flash?

Inkling-Small is the stronger model overall, scoring 46.5 to 42.5 on the Noometry Index. Qwen3.5-Flash costs 3.6× less per token, which makes it the better buy when Inkling-Small's lead doesn't matter for your workload.

Which is cheaper, Inkling-Small or Qwen3.5-Flash?

Qwen3.5-Flash is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; Inkling-Small lists at $0.45 and $1.20.

Is Inkling-Small or Qwen3.5-Flash better for coding?

Inkling-Small scores higher on coding benchmarks: 43.6 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

Qwen3.5-Flash does, with 1M tokens against 524K.

How many benchmarks do Inkling-Small and Qwen3.5-Flash share?

25 benchmarks have published results for both models. Inkling-Small has 33 scored results on Noometry and Qwen3.5-Flash has 32.

Related comparisons

Go deeper