Thinking Machines Lab, open weights
Inkling
Inkling by Thinking Machines Lab ranks 80th of 354 ranked models on the Noometry Index as of October 2026, with a score of 44.1. Its strongest category is knowledge, where it ranks 49th. API pricing starts at $1.87 per million input tokens and $4.68 per million output tokens, with a 66K-token context window.
Last verified
Specifications
- Noometry rank
- #80 of 354
- Index score
- 44.1
- Evidence
- Confirmed 41 results
- Provider
- Thinking Machines Lab
- Released
- July 15, 2026
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 66K
- Max output
- 66K
- Input price
- $1.87 / M
- Output price
- $4.68 / M
- Blended price
- $2.57 / M
- Output speed
- Not measured
- Value
- #159 of 219
- Knowledge cutoff
- Unknown
- Input
- text, image
- Hugging Face
- thinkingmachines/Inkling
Category scores
Each category score combines every public result we have in that category.
- Coding 34.5
- Agentic & Tool Use 29.6
- Reasoning 40.4
- Math 31.3
- Knowledge 55.1
- Multilingual 54.0
- Instruction Following 75.1
- Long Context 43.8
- Writing & Preference 65.2
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 34.5 | #234 | 6 |
| Agentic & Tool Use | 29.6 | #85 | 2 |
| Reasoning | 40.4 | #56 | 8 |
| Math | 31.3 | #225 | 5 |
| Knowledge | 55.1 | #49 | 3 |
| Multilingual | 54.0 | #52 | 1 |
| Instruction Following | 75.1 | #71 | 1 |
| Long Context | 43.8 | #86 | 1 |
| Writing & Preference | 65.2 | #51 | 5 |
Strengths and weaknesses
Categories where Inkling places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Knowledge | 55.1 | +17.8 | #49 of 314, top 16% |
| Reasoning | 40.4 | +16.8 | #56 of 350, top 16% |
| Writing & Preference | 65.2 | +11.4 | #51 of 312, top 17% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Coding | 34.5 | −4.2 | #234 of 340, top 69% |
| Math | 31.3 | −5.2 | #225 of 327, top 69% |
| Agentic & Tool Use | 29.6 | −0.8 | #85 of 154, top 56% |
Closest competitors
The models ranked just above and below Inkling. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| GPT-5.4 mini | #76 | 45.0 | $1.69 | 10 | Compare |
| Amazon Nova Experimental Chat 26 02 10 | #77 | 44.5 | — | — | Compare |
| DeepSeek-V3.2-Exp | #78 | 44.3 | $0.29 | 16 | Compare |
| Hy3 | #79 | 44.2 | $0.14 | — | Compare |
| Claude Sonnet 4.5 | #81 | 44.1 | $6 | 85 | Compare |
| Chatgpt 4o Latest 20250326 | #82 | 43.8 | — | 21 | Compare |
| ERNIE 5.1 | #83 | 43.8 | — | — | Compare |
| GLM-5V-Turbo | #84 | 43.8 | $1.90 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierCode | 14% | #34 of 37, top 92% | Epoch AI | ||
| LMArena WebDev | 1413 | #67 of 113, top 60% | LMArena | 2026-10-08 | |
| FrontierSWE | 4.1% | #18 of 18, top 100% | xhigh | Epoch AI | |
| SciCode | 47% | #53 of 121, top 44% | xhigh | Epoch AI | |
| WeirdML | 32.3% | #95 of 119, top 80% | high | Epoch AI | |
| LMArena Coding | 1464 | #67 of 294, top 23% | LMArena | 2026-10-08 | |
| ALE-Bench | 946 | #44 of 105, top 42% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| APEX-Agents | 33.8% | #42 of 49, top 86% | Epoch AI | ||
| τ²-bench Banking | 25% | #18 of 26, top 70% | max | τ²-bench | 2026-08-04 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 36.5% | #38 of 83, top 46% | Epoch AI | ||
| SimpleBench | 50% | #40 of 77, top 52% | Epoch AI | ||
| ARC-AGI-1 | 79.5% | #39 of 83, top 47% | Epoch AI | ||
| CritPt | 5.4% | #56 of 134, top 42% | xhigh | Epoch AI | |
| Chess Puzzles | 21% | #54 of 129, top 42% | xhigh | Epoch AI | 2026-08-05 |
| LMArena Hard Prompts | 1451 | #61 of 297, top 21% | LMArena | 2026-10-08 | |
| DTBench | 87.5% | #50 of 151, top 34% | xhigh | Epoch AI | |
| LMCA | 37.6% | #54 of 125, top 44% | xhigh | Epoch AI | |
| Epoch Capabilities Index | 148.54 | #61 of 213, top 29% | Epoch AI | 2026-07-15 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 33.3% | #59 of 81, top 73% | xhigh | Epoch AI | 2026-08-06 |
| FrontierMath Tier 4 | 4.9% | #55 of 63, top 88% | xhigh | Epoch AI | 2026-08-06 |
| OTIS Mock AIME 2024-2025 | 88.9% | #56 of 173, top 33% | xhigh | Epoch AI | 2026-08-05 |
| ProofBench | 0% | #75 of 77, top 98% | Epoch AI | ||
| LMArena Math | 1479 | #29 of 285, top 11% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 88.3% | #48 of 186, top 26% | xhigh | Epoch AI | 2026-08-05 |
| SimpleQA Verified | 40.3% | #44 of 77, top 58% | xhigh | Epoch AI | 2026-08-27 |
| LMArena Expert | 1465 | #55 of 273, top 21% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1434 | #52 of 297, top 18% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1490 | #56 of 285, top 20% | LMArena | 2026-10-08 | |
| LMArena French | 1458 | #55 of 223, top 25% | LMArena | 2026-10-08 | |
| LMArena German | 1446 | #50 of 231, top 22% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1429 | #33 of 211, top 16% | LMArena | 2026-10-08 | |
| LMArena Korean | 1404 | #50 of 213, top 24% | LMArena | 2026-10-08 | |
| LMArena Russian | 1429 | #68 of 283, top 25% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1448 | #56 of 226, top 25% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1426 | #66 of 298, top 23% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1434 | #75 of 291, top 26% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1441 | #59 of 297, top 20% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1387 | #97 of 295, top 33% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 1611 | #40 of 115, top 35% | EQ-Bench | ||
| EQ-Bench 4 | 1226 | #12 of 28, top 43% | EQ-Bench | ||
| LMArena Multi-Turn | 1436 | #76 of 295, top 26% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| deepinfra | $0.95 | $4.05 | $0.16 | 2026-10-10 |
| fireworks | $1 | $4.05 | $0.17 | 2026-10-10 |
| openrouter | $1 | $4.05 | $0.17 | 2026-10-10 |
| thinking-machines | $1.87 | $4.68 | $0.37 | 2026-10-10 |
| together | $1 | $4.05 | $0.17 | 2026-10-10 |
Compare Inkling
- Inkling vs Hy3
- Inkling vs Claude Sonnet 4.5
- Inkling vs DeepSeek-V3.2-Exp
- Inkling vs Chatgpt 4o Latest 20250326
- Inkling vs Amazon Nova Experimental Chat 26 02 10
- Inkling vs ERNIE 5.1
- Inkling vs GPT-6 Astra
- Inkling vs Claude Fable 5.1
- Inkling vs Gemini 3.8 Flash
- Inkling vs Kimi K3
- Inkling vs Grok 4.6
- Inkling vs Qwen3.8 Max
- Inkling vs GLM-5.3
- Inkling vs Muse Spark 1.3
Other Thinking Machines Lab models
- Inkling-Small46.5
Frequently asked questions
How good is Inkling?
Inkling by Thinking Machines Lab ranks 80th of 354 ranked models on the Noometry Index as of October 2026, with a score of 44.1. Its strongest category is knowledge, where it ranks 49th. API pricing starts at $1.87 per million input tokens and $4.68 per million output tokens, with a 66K-token context window.
How much does Inkling cost?
Inkling costs $1.87 per million input tokens and $4.68 per million output tokens on Thinking Machines Lab's own API, with cached input at $0.37.
What is Inkling's context window?
Inkling accepts up to 66K tokens of input and can write up to 66K tokens in one response.
Is Inkling open source?
Yes. Inkling's weights are downloadable from Hugging Face (thinkingmachines/Inkling); check the license for commercial terms.
What are Inkling's strengths and weaknesses?
Relative to other ranked models, Inkling places best in knowledge, reasoning, writing & preference and lowest in coding, math, agentic & tool use.
What is Inkling best at?
Its best category is knowledge, where it ranks 49th on Noometry.