Model comparison

Inkling vs Qwen3 32B

Inkling is the stronger model overall, scoring 44.1 to 39.2 on the Noometry Index. Qwen3 32B costs 2.1× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Last verified . 21 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Inkling scores higher in 6 categories and Qwen3 32B in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Inkling leads 40.4 to 20.2.
  • The biggest single-benchmark swing is GPQA Diamond: 88.3% for Inkling and 65.7% for Qwen3 32B.
  • Qwen3 32B is cheaper at $0.70 / $2.80 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • Qwen3 32B accepts more context: 131K tokens versus 66K.

Side by side

Inkling and Qwen3 32B specifications
InklingQwen3 32B
ProviderThinking Machines LabAlibaba (Qwen)
Noometry Index44.139.2
Released2026-07-152025-04
WeightsOpenOpen
Context window66K131K
Max output66K16K
Input $ / M tokens$1.87$0.70
Output $ / M tokens$4.68$2.80
Results tracked4126

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 32B leads

Inkling: 34.5 (#234), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkInklingQwen3 32B
SciCode47%35.4%
LMArena Coding14641358
FrontierCode14%—
Aider Polyglot—40%
LMArena WebDev1413—
FrontierSWE4.1%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Qwen3 32B leads

Inkling: 29.6 (#85), Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkInklingQwen3 32B
APEX-Agents33.8%—
Berkeley Function Calling Leaderboard—48.7%
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkInklingQwen3 32B
CritPt5.4%0.3%
Chess Puzzles21%5%
LMArena Hard Prompts14511334
DTBench87.5%67.5%
LMCA37.6%17.3%
Epoch Capabilities Index148.54138.51
ARC-AGI-236.5%—
SimpleBench50%—
Kagi LLM Benchmark—54.9%
ARC-AGI-179.5%—

Math Qwen3 32B leads

Inkling: 31.3 (#225), Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkInklingQwen3 32B
OTIS Mock AIME 2024-202588.9%66.9%
LMArena Math14791399
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkInklingQwen3 32B
GPQA Diamond88.3%65.7%
LMArena Expert14651362
SimpleQA Verified40.3%—
Vectara Hallucination Rate—5.9%

Multilingual Inkling leads

Inkling: 54.0 (#52), Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkInklingQwen3 32B
LMArena Non-English14341317
LMArena Chinese14901357
LMArena German14461341
LMArena Russian14291311
LMArena French1458—
LMArena Japanese1429—
LMArena Korean1404—
LMArena Spanish1448—

Instruction Following Inkling leads

Inkling: 75.1 (#71), Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkInklingQwen3 32B
LMArena Instruction Following14261305

Long Context Too close to call

Inkling: 43.8 (#86), Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkInklingQwen3 32B
LMArena Longer Query14341327
Fiction.LiveBench—74.2%

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkInklingQwen3 32B
LMArena Text14411340
LMArena Creative Writing13871297
LMArena Multi-Turn14361331
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Qwen3 32B?

Inkling is the stronger model overall, scoring 44.1 to 39.2 on the Noometry Index. Qwen3 32B costs 2.1× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Which is cheaper, Inkling or Qwen3 32B?

Qwen3 32B is cheaper. It lists at $0.70 per million input tokens and $2.80 per million output tokens; Inkling lists at $1.87 and $4.68.

Is Inkling or Qwen3 32B better for coding?

Qwen3 32B scores higher on coding benchmarks: 37.7 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

Qwen3 32B does, with 131K tokens against 66K.

How many benchmarks do Inkling and Qwen3 32B share?

21 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper