Model comparison

Inkling-Small vs Qwen3.5 35B-A3B

Inkling-Small is the stronger model overall, scoring 46.5 to 42.0 on the Noometry Index.

Last verified . 24 shared benchmarks.

Inkling-Small Thinking Machines Lab

46.5

Rank #63 Confirmed

Qwen3.5 35B-A3B Alibaba (Qwen)

42.0

Rank #123 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Inkling-Small scores higher in 8 categories and Qwen3.5 35B-A3B in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Inkling-Small leads 38.6 to 24.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 90% for Inkling-Small and 70% for Qwen3.5 35B-A3B.
  • Inkling-Small is cheaper at $0.45 / $1.20 per million input/output tokens, against $0.25 / $2 for Qwen3.5 35B-A3B.
  • Inkling-Small accepts more context: 524K tokens versus 262K.

Side by side

Inkling-Small and Qwen3.5 35B-A3B specifications
Inkling-SmallQwen3.5 35B-A3B
ProviderThinking Machines LabAlibaba (Qwen)
Noometry Index46.542.0
Released2026-07-152026-02-01
WeightsOpenOpen
Context window524K262K
Max output1.05M66K
Input $ / M tokens$0.45$0.25
Output $ / M tokens$1.20$2
Results tracked3328

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling-Small leads

Inkling-Small: 43.6 (#85), Qwen3.5 35B-A3B: 33.8 (#251)

Coding benchmarks
BenchmarkInkling-SmallQwen3.5 35B-A3B
LMArena WebDev14091254
SciCode48.7%29.3%
LMArena Coding14511410

Reasoning Inkling-Small leads

Inkling-Small: 38.6 (#63), Qwen3.5 35B-A3B: 24.6 (#161)

Reasoning benchmarks
BenchmarkInkling-SmallQwen3.5 35B-A3B
CritPt8.3%0.6%
Chess Puzzles18%10%
LMArena Hard Prompts14231400
Epoch Capabilities Index150.15142.52
ARC-AGI-240.1%—
ARC-AGI-184%—
Mystery Game Puzzles6%—
DTBench—80%
LMCA—29.5%

Math Inkling-Small leads

Inkling-Small: 45.1 (#77), Qwen3.5 35B-A3B: 39.9 (#97)

Math benchmarks
BenchmarkInkling-SmallQwen3.5 35B-A3B
OTIS Mock AIME 2024-202590%70%
LMArena Math14591404
FrontierMath (Tiers 1-3)46.3%—
FrontierMath Tier 417.1%—
MathArena Final-Answer Competitions—56%
ProofBench6%—

Knowledge Too close to call

Inkling-Small: 48.2 (#77), Qwen3.5 35B-A3B: 47.8 (#79)

Knowledge benchmarks
BenchmarkInkling-SmallQwen3.5 35B-A3B
GPQA Diamond88.5%83.5%
LMArena Expert14421408
SimpleQA Verified19.1%—
Vectara Hallucination Rate—10.5%

Multimodal Not comparable

Inkling-Small: 39.1 (#62), Qwen3.5 35B-A3B: —

Multimodal benchmarks
BenchmarkInkling-SmallQwen3.5 35B-A3B
LMArena Vision1235—

Multilingual Inkling-Small leads

Inkling-Small: 51.7 (#104), Qwen3.5 35B-A3B: 50.0 (#127)

Multilingual benchmarks
BenchmarkInkling-SmallQwen3.5 35B-A3B
LMArena Non-English14021378
LMArena Chinese14651457
LMArena French14361412
LMArena German14051367
LMArena Japanese14051325
LMArena Korean13631356
LMArena Russian13911376
LMArena Spanish14281392

Instruction Following Inkling-Small leads

Inkling-Small: 73.8 (#114), Qwen3.5 35B-A3B: 72.8 (#128)

Instruction Following benchmarks
BenchmarkInkling-SmallQwen3.5 35B-A3B
LMArena Instruction Following13991379

Long Context Too close to call

Inkling-Small: 42.7 (#118), Qwen3.5 35B-A3B: 42.4 (#127)

Long Context benchmarks
BenchmarkInkling-SmallQwen3.5 35B-A3B
LMArena Longer Query14011389

Writing & Preference Inkling-Small leads

Inkling-Small: 59.6 (#107), Qwen3.5 35B-A3B: 57.9 (#124)

Writing & Preference benchmarks
BenchmarkInkling-SmallQwen3.5 35B-A3B
LMArena Text14141395
LMArena Creative Writing13311346
LMArena Multi-Turn14181390
EQ-Bench Creative Writing1491—

Frequently asked questions

Is Inkling-Small better than Qwen3.5 35B-A3B?

Inkling-Small is the stronger model overall, scoring 46.5 to 42.0 on the Noometry Index.

Which is cheaper, Inkling-Small or Qwen3.5 35B-A3B?

Inkling-Small is cheaper. It lists at $0.45 per million input tokens and $1.20 per million output tokens; Qwen3.5 35B-A3B lists at $0.25 and $2.

Is Inkling-Small or Qwen3.5 35B-A3B better for coding?

Inkling-Small scores higher on coding benchmarks: 43.6 versus 33.8 in the Noometry coding category.

Which has the bigger context window?

Inkling-Small does, with 524K tokens against 262K.

How many benchmarks do Inkling-Small and Qwen3.5 35B-A3B share?

24 benchmarks have published results for both models. Inkling-Small has 33 scored results on Noometry and Qwen3.5 35B-A3B has 28.

Related comparisons

Go deeper