Model comparison

Inkling vs Qwen3.5 35B-A3B

Inkling is the stronger model overall, scoring 44.1 to 42.0 on the Noometry Index. Qwen3.5 35B-A3B costs 3.7× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Last verified . 26 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Qwen3.5 35B-A3B Alibaba (Qwen)

42.0

Rank #123 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Inkling scores higher in 7 categories and Qwen3.5 35B-A3B in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Inkling leads 40.4 to 24.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 88.9% for Inkling and 70% for Qwen3.5 35B-A3B.
  • Qwen3.5 35B-A3B is cheaper at $0.25 / $2 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • Qwen3.5 35B-A3B accepts more context: 262K tokens versus 66K.

Side by side

Inkling and Qwen3.5 35B-A3B specifications
InklingQwen3.5 35B-A3B
ProviderThinking Machines LabAlibaba (Qwen)
Noometry Index44.142.0
Released2026-07-152026-02-01
WeightsOpenOpen
Context window66K262K
Max output66K66K
Input $ / M tokens$1.87$0.25
Output $ / M tokens$4.68$2
Results tracked4128

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Inkling: 34.5 (#234), Qwen3.5 35B-A3B: 33.8 (#251)

Coding benchmarks
BenchmarkInklingQwen3.5 35B-A3B
LMArena WebDev14131254
SciCode47%29.3%
LMArena Coding14641410
FrontierCode14%—
FrontierSWE4.1%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Qwen3.5 35B-A3B: —

Agentic & Tool Use benchmarks
BenchmarkInklingQwen3.5 35B-A3B
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Qwen3.5 35B-A3B: 24.6 (#161)

Reasoning benchmarks
BenchmarkInklingQwen3.5 35B-A3B
CritPt5.4%0.6%
Chess Puzzles21%10%
LMArena Hard Prompts14511400
DTBench87.5%80%
LMCA37.6%29.5%
Epoch Capabilities Index148.54142.52
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—

Math Qwen3.5 35B-A3B leads

Inkling: 31.3 (#225), Qwen3.5 35B-A3B: 39.9 (#97)

Math benchmarks
BenchmarkInklingQwen3.5 35B-A3B
OTIS Mock AIME 2024-202588.9%70%
LMArena Math14791404
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
MathArena Final-Answer Competitions—56%
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Qwen3.5 35B-A3B: 47.8 (#79)

Knowledge benchmarks
BenchmarkInklingQwen3.5 35B-A3B
GPQA Diamond88.3%83.5%
LMArena Expert14651408
SimpleQA Verified40.3%—
Vectara Hallucination Rate—10.5%

Multilingual Inkling leads

Inkling: 54.0 (#52), Qwen3.5 35B-A3B: 50.0 (#127)

Multilingual benchmarks
BenchmarkInklingQwen3.5 35B-A3B
LMArena Non-English14341378
LMArena Chinese14901457
LMArena French14581412
LMArena German14461367
LMArena Japanese14291325
LMArena Korean14041356
LMArena Russian14291376
LMArena Spanish14481392

Instruction Following Inkling leads

Inkling: 75.1 (#71), Qwen3.5 35B-A3B: 72.8 (#128)

Instruction Following benchmarks
BenchmarkInklingQwen3.5 35B-A3B
LMArena Instruction Following14261379

Long Context Inkling leads

Inkling: 43.8 (#86), Qwen3.5 35B-A3B: 42.4 (#127)

Long Context benchmarks
BenchmarkInklingQwen3.5 35B-A3B
LMArena Longer Query14341389

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Qwen3.5 35B-A3B: 57.9 (#124)

Writing & Preference benchmarks
BenchmarkInklingQwen3.5 35B-A3B
LMArena Text14411395
LMArena Creative Writing13871346
LMArena Multi-Turn14361390
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Qwen3.5 35B-A3B?

Inkling is the stronger model overall, scoring 44.1 to 42.0 on the Noometry Index. Qwen3.5 35B-A3B costs 3.7× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Which is cheaper, Inkling or Qwen3.5 35B-A3B?

Qwen3.5 35B-A3B is cheaper. It lists at $0.25 per million input tokens and $2 per million output tokens; Inkling lists at $1.87 and $4.68.

Is Inkling or Qwen3.5 35B-A3B better for coding?

They score almost the same on coding (34.5 vs 33.8); test both on your own repository before choosing.

Which has the bigger context window?

Qwen3.5 35B-A3B does, with 262K tokens against 66K.

How many benchmarks do Inkling and Qwen3.5 35B-A3B share?

26 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Qwen3.5 35B-A3B has 28.

Related comparisons

Go deeper