Model comparison

Inkling vs Mistral Medium 3.5

Inkling is the stronger model overall, scoring 44.1 to 40.2 on the Noometry Index.

Last verified . 19 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Inkling scores higher in 6 categories and Mistral Medium 3.5 in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Inkling leads 40.4 to 17.3.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium 3.5.
  • Mistral Medium 3.5 accepts more context: 262K tokens versus 66K.

Side by side

Inkling and Mistral Medium 3.5 specifications
InklingMistral Medium 3.5
ProviderThinking Machines LabMistral AI
Noometry Index44.140.2
Released2026-07-15—
WeightsOpenOpen
Context window66K262K
Max output66K210K
Input $ / M tokens$1.87$1.50
Output $ / M tokens$4.68$7.50
Results tracked4122

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium 3.5 leads

Inkling: 34.5 (#234), Mistral Medium 3.5: 36.0 (#213)

Coding benchmarks
BenchmarkInklingMistral Medium 3.5
LMArena WebDev14131264
LMArena Coding14641461
FrontierCode14%—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Mistral Medium 3.5: —

Agentic & Tool Use benchmarks
BenchmarkInklingMistral Medium 3.5
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Mistral Medium 3.5: 17.3 (#295)

Reasoning benchmarks
BenchmarkInklingMistral Medium 3.5
LMArena Hard Prompts14511436
Epoch Capabilities Index148.54141.35
ARC-AGI-236.5%—
SimpleBench50%—
Kagi LLM Benchmark—41.4%
NYT Connections (extended)—12.9%
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
DTBench87.5%—
LMCA37.6%—

Math Mistral Medium 3.5 leads

Inkling: 31.3 (#225), Mistral Medium 3.5: 39.1 (#113)

Math benchmarks
BenchmarkInklingMistral Medium 3.5
LMArena Math14791431
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—

Knowledge Inkling leads

Inkling: 55.1 (#49), Mistral Medium 3.5: 40.0 (#126)

Knowledge benchmarks
BenchmarkInklingMistral Medium 3.5
LMArena Expert14651432
GPQA Diamond88.3%—
SimpleQA Verified40.3%—

Multimodal Not comparable

Inkling: —, Mistral Medium 3.5: 38.3 (#65)

Multimodal benchmarks
BenchmarkInklingMistral Medium 3.5
LMArena Vision—1223

Multilingual Inkling leads

Inkling: 54.0 (#52), Mistral Medium 3.5: 51.9 (#100)

Multilingual benchmarks
BenchmarkInklingMistral Medium 3.5
LMArena Non-English14341404
LMArena Chinese14901442
LMArena French14581448
LMArena German14461451
LMArena Korean14041385
LMArena Russian14291395
LMArena Spanish14481409
LMArena Japanese1429—

Instruction Following Too close to call

Inkling: 75.1 (#71), Mistral Medium 3.5: 74.6 (#90)

Instruction Following benchmarks
BenchmarkInklingMistral Medium 3.5
LMArena Instruction Following14261415

Long Context Too close to call

Inkling: 43.8 (#86), Mistral Medium 3.5: 43.2 (#103)

Long Context benchmarks
BenchmarkInklingMistral Medium 3.5
LMArena Longer Query14341415

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Mistral Medium 3.5: 58.5 (#117)

Writing & Preference benchmarks
BenchmarkInklingMistral Medium 3.5
LMArena Text14411421
LMArena Creative Writing13871374
EQ-Bench 41226993
LMArena Multi-Turn14361423
EQ-Bench Creative Writing1611—

Frequently asked questions

Is Inkling better than Mistral Medium 3.5?

Inkling is the stronger model overall, scoring 44.1 to 40.2 on the Noometry Index.

Which is cheaper, Inkling or Mistral Medium 3.5?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; Mistral Medium 3.5 lists at $1.50 and $7.50.

Is Inkling or Mistral Medium 3.5 better for coding?

Mistral Medium 3.5 scores higher on coding benchmarks: 36.0 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium 3.5 does, with 262K tokens against 66K.

How many benchmarks do Inkling and Mistral Medium 3.5 share?

19 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Mistral Medium 3.5 has 22.

Related comparisons

Go deeper