Model comparison

Inkling vs Mistral Small

Inkling is the stronger model overall, scoring 44.1 to 33.4 on the Noometry Index. Mistral Small costs 9.8× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Last verified . 24 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Inkling scores higher in 9 categories and Mistral Small in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 31.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 88.9% for Inkling and 5.8% for Mistral Small.
  • Mistral Small is cheaper at $0.15 / $0.60 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • Mistral Small accepts more context: 262K tokens versus 66K.

Side by side

Inkling and Mistral Small specifications
InklingMistral Small
ProviderThinking Machines LabMistral AI
Noometry Index44.133.4
Released2026-07-152024-02-26
WeightsOpenOpen
Context window66K262K
Max output66K256K
Input $ / M tokens$1.87$0.15
Output $ / M tokens$4.68$0.60
Results tracked4139

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Inkling: 34.5 (#234), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkInklingMistral Small
SciCode47%26.5%
LMArena Coding14641362
ALE-Bench946497.62
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
WeirdML32.3%—
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
BigCodeBench Complete—46.6%

Agentic & Tool Use Inkling leads

Inkling: 29.6 (#85), Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkInklingMistral Small
APEX-Agents33.8%—
Berkeley Function Calling Leaderboard—37.1%
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkInklingMistral Small
CritPt5.4%0%
LMArena Hard Prompts14511335
DTBench87.5%70.9%
LMCA37.6%20.6%
ARC-AGI-236.5%—
SimpleBench50%—
Kagi LLM Benchmark—37.8%
ARC-AGI-179.5%—
Chess Puzzles21%—
LiveBench Reasoning—44.8%
LiveBench Data Analysis—53.7%
Epoch Capabilities Index148.54—
LiveBench—44%

Math Inkling leads

Inkling: 31.3 (#225), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkInklingMistral Small
OTIS Mock AIME 2024-202588.9%5.8%
LMArena Math14791341
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
ProofBench0%—
LiveBench Math—39.9%
MATH Level 5—46.8%

Knowledge Inkling leads

Inkling: 55.1 (#49), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkInklingMistral Small
GPQA Diamond88.3%47.5%
LMArena Expert14651291
SimpleQA Verified40.3%—
Vectara Hallucination Rate—5.1%
MMLU—68.7%

Multimodal Not comparable

Inkling: —, Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkInklingMistral Small
LMArena Vision—1142

Multilingual Inkling leads

Inkling: 54.0 (#52), Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkInklingMistral Small
LMArena Non-English14341315
LMArena Chinese14901340
LMArena French14581337
LMArena German14461340
LMArena Japanese14291275
LMArena Korean14041259
LMArena Russian14291324
LMArena Spanish14481346

Instruction Following Inkling leads

Inkling: 75.1 (#71), Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkInklingMistral Small
LMArena Instruction Following14261310
LiveBench Instruction Following—63.7%

Long Context Inkling leads

Inkling: 43.8 (#86), Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkInklingMistral Small
LMArena Longer Query14341327

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkInklingMistral Small
LMArena Text14411338
LMArena Creative Writing13871305
LMArena Multi-Turn14361344
EQ-Bench Creative Writing1611—
EQ-Bench 41226—
LiveBench Language—30.5%

Frequently asked questions

Is Inkling better than Mistral Small?

Inkling is the stronger model overall, scoring 44.1 to 33.4 on the Noometry Index. Mistral Small costs 9.8× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Which is cheaper, Inkling or Mistral Small?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Inkling lists at $1.87 and $4.68.

Is Inkling or Mistral Small better for coding?

They score almost the same on coding (34.5 vs 34.0); test both on your own repository before choosing.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 66K.

How many benchmarks do Inkling and Mistral Small share?

24 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper