Model comparison

Inkling vs Mistral Large

Inkling is the stronger model overall, scoring 44.1 to 31.9 on the Noometry Index.

Last verified . 27 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Inkling scores higher in 9 categories and Mistral Large in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 30.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 88.9% for Inkling and 8.5% for Mistral Large.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Mistral Large accepts more context: 131K tokens versus 66K.

Side by side

Inkling and Mistral Large specifications
InklingMistral Large
ProviderThinking Machines LabMistral AI
Noometry Index44.131.9
Released2026-07-152024-02-26
WeightsOpenOpen
Context window66K131K
Max output66K16K
Input $ / M tokens$1.87$2
Output $ / M tokens$4.68$6
Results tracked4151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Inkling: 34.5 (#234), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkInklingMistral Large
SciCode47%36.2%
LMArena Coding14641277
ALE-Bench946264.7
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
WeirdML32.3%—
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Too close to call

Inkling: 29.6 (#85), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkInklingMistral Large
APEX-Agents33.8%—
Berkeley Function Calling Leaderboard—38.4%
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkInklingMistral Large
SimpleBench50%22.5%
CritPt5.4%0%
LMArena Hard Prompts14511257
DTBench87.5%65.1%
LMCA37.6%16.7%
Epoch Capabilities Index148.54128.52
ARC-AGI-236.5%—
ARC-AGI-179.5%—
Chess Puzzles21%—
LiveBench Reasoning—43.5%
LiveBench Data Analysis—50.1%
ForecastBench—57.1
LiveBench—48.4%

Math Inkling leads

Inkling: 31.3 (#225), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkInklingMistral Large
OTIS Mock AIME 2024-202588.9%8.5%
LMArena Math14791262
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
ProofBench0%—
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Inkling leads

Inkling: 55.1 (#49), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkInklingMistral Large
GPQA Diamond88.3%51.3%
LMArena Expert14651232
SimpleQA Verified40.3%—
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
MMLU—80%

Multilingual Inkling leads

Inkling: 54.0 (#52), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkInklingMistral Large
LMArena Non-English14341237
LMArena Chinese14901240
LMArena French14581325
LMArena German14461254
LMArena Japanese14291188
LMArena Korean14041202
LMArena Russian14291257
LMArena Spanish14481268

Instruction Following Inkling leads

Inkling: 75.1 (#71), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkInklingMistral Large
LMArena Instruction Following14261249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Inkling leads

Inkling: 43.8 (#86), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkInklingMistral Large
LMArena Longer Query14341261

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkInklingMistral Large
LMArena Text14411266
LMArena Creative Writing13871243
EQ-Bench Creative Writing1611985
LMArena Multi-Turn14361260
Short-Story Creative Writing—69%
WildBench—80.1%
EQ-Bench 41226—
LiveBench Language—39.4%

Frequently asked questions

Is Inkling better than Mistral Large?

Inkling is the stronger model overall, scoring 44.1 to 31.9 on the Noometry Index.

Which is cheaper, Inkling or Mistral Large?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; Mistral Large lists at $2 and $6.

Is Inkling or Mistral Large better for coding?

They score almost the same on coding (34.5 vs 34.3); test both on your own repository before choosing.

Which has the bigger context window?

Mistral Large does, with 131K tokens against 66K.

How many benchmarks do Inkling and Mistral Large share?

27 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper