Model comparison

Inkling vs Qwen1.5 4b Chat

Inkling is the stronger model overall, scoring 44.1 to 28.8 on the Noometry Index.

Last verified . 13 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Inkling scores higher in 8 categories and Qwen1.5 4b Chat in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Inkling leads 65.2 to 23.8.

Side by side

Inkling and Qwen1.5 4b Chat specifications
InklingQwen1.5 4b Chat
ProviderThinking Machines LabAlibaba (Qwen)
Noometry Index44.128.8
Released2026-07-15—
WeightsOpenOpen
Context window66K—
Max output66K—
Input $ / M tokens$1.87—
Output $ / M tokens$4.68—
Results tracked4113

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling leads

Inkling: 34.5 (#234), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkInklingQwen1.5 4b Chat
LMArena Coding1464999
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Qwen1.5 4b Chat: —

Agentic & Tool Use benchmarks
BenchmarkInklingQwen1.5 4b Chat
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkInklingQwen1.5 4b Chat
LMArena Hard Prompts1451976
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Too close to call

Inkling: 31.3 (#225), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkInklingQwen1.5 4b Chat
LMArena Math14791026
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkInklingQwen1.5 4b Chat
LMArena Expert1465980
GPQA Diamond88.3%—
SimpleQA Verified40.3%—

Multilingual Inkling leads

Inkling: 54.0 (#52), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkInklingQwen1.5 4b Chat
LMArena Non-English1434979
LMArena Chinese14901024
LMArena German1446902
LMArena Russian1429952
LMArena French1458—
LMArena Japanese1429—
LMArena Korean1404—
LMArena Spanish1448—

Instruction Following Inkling leads

Inkling: 75.1 (#71), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkInklingQwen1.5 4b Chat
LMArena Instruction Following1426978

Long Context Inkling leads

Inkling: 43.8 (#86), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkInklingQwen1.5 4b Chat
LMArena Longer Query1434988

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkInklingQwen1.5 4b Chat
LMArena Text1441997
LMArena Creative Writing1387969
LMArena Multi-Turn1436977
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Qwen1.5 4b Chat?

Inkling is the stronger model overall, scoring 44.1 to 28.8 on the Noometry Index.

Is Inkling or Qwen1.5 4b Chat better for coding?

Inkling scores higher on coding benchmarks: 34.5 versus 29.1 in the Noometry coding category.

How many benchmarks do Inkling and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper