Model comparison

Inkling vs Mistral Medium

Inkling is the stronger model overall, scoring 44.1 to 36.3 on the Noometry Index.

Last verified . 27 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Inkling scores higher in 9 categories and Mistral Medium in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 25.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 88.9% for Inkling and 32.2% for Mistral Medium.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium.
  • Mistral Medium accepts more context: 262K tokens versus 66K.

Side by side

Inkling and Mistral Medium specifications
InklingMistral Medium
ProviderThinking Machines LabMistral AI
Noometry Index44.136.3
Released2026-07-152023-12-11
WeightsOpenOpen
Context window66K262K
Max output66K262K
Input $ / M tokens$1.87$1.50
Output $ / M tokens$4.68$7.50
Results tracked4136

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Inkling: 34.5 (#234), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkInklingMistral Medium
FrontierCode14%8%
SciCode47%40.2%
WeirdML32.3%43.7%
LMArena Coding14641434
ALE-Bench946763.98
LMArena WebDev1413—
FrontierSWE4.1%—

Agentic & Tool Use Inkling leads

Inkling: 29.6 (#85), Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkInklingMistral Medium
APEX-Agents33.8%—
Berkeley Function Calling Leaderboard—37.7%
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkInklingMistral Medium
CritPt5.4%0%
LMArena Hard Prompts14511426
DTBench87.5%75.5%
LMCA37.6%26.1%
ARC-AGI-236.5%—
SimpleBench50%—
Kagi LLM Benchmark—50%
ARC-AGI-179.5%—
Chess Puzzles21%—
Surface Evolver Bench—26.9%
Epoch Capabilities Index148.54—

Math Inkling leads

Inkling: 31.3 (#225), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkInklingMistral Medium
OTIS Mock AIME 2024-202588.9%32.2%
ProofBench0%9%
LMArena Math14791408
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Inkling leads

Inkling: 55.1 (#49), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkInklingMistral Medium
GPQA Diamond88.3%59.5%
LMArena Expert14651408
Humanity's Last Exam—4.5%
SimpleQA Verified40.3%—
Vectara Hallucination Rate—22.7%

Multimodal Not comparable

Inkling: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkInklingMistral Medium
LMArena Vision—1172

Multilingual Inkling leads

Inkling: 54.0 (#52), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkInklingMistral Medium
LMArena Non-English14341408
LMArena Chinese14901447
LMArena French14581459
LMArena German14461432
LMArena Japanese14291378
LMArena Korean14041380
LMArena Russian14291411
LMArena Spanish14481433

Instruction Following Inkling leads

Inkling: 75.1 (#71), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkInklingMistral Medium
LMArena Instruction Following14261398

Long Context Too close to call

Inkling: 43.8 (#86), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkInklingMistral Medium
LMArena Longer Query14341406

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkInklingMistral Medium
LMArena Text14411424
LMArena Creative Writing13871391
LMArena Multi-Turn14361418
Short-Story Creative Writing—77.3%
EQ-Bench Creative Writing1611—
EQ-Bench 41226—

Frequently asked questions

Is Inkling better than Mistral Medium?

Inkling is the stronger model overall, scoring 44.1 to 36.3 on the Noometry Index.

Which is cheaper, Inkling or Mistral Medium?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Is Inkling or Mistral Medium better for coding?

They score almost the same on coding (34.5 vs 34.2); test both on your own repository before choosing.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 66K.

How many benchmarks do Inkling and Mistral Medium share?

27 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper