Model comparison

Inkling vs Pixtral Large

Inkling is the stronger model overall, scoring 44.1 to 32.2 on the Noometry Index.

Last verified . 1 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Summary

  • They share 1 benchmark with published results for both. Inkling scores higher in 2 categories and Pixtral Large in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Inkling leads 65.2 to 32.9.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $2 / $6 for Pixtral Large.
  • Pixtral Large accepts more context: 128K tokens versus 66K.

Side by side

Inkling and Pixtral Large specifications
InklingPixtral Large
ProviderThinking Machines LabMistral AI
Noometry Index44.132.2
Released2026-07-152024-11-01
WeightsOpenOpen
Context window66K128K
Max output66K128K
Input $ / M tokens$1.87$2
Output $ / M tokens$4.68$6
Results tracked413

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Inkling: 34.5 (#234), Pixtral Large: —

Coding benchmarks
BenchmarkInklingPixtral Large
FrontierCode14%—
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
WeirdML32.3%—
LMArena Coding1464—
ALE-Bench946—

Agentic & Tool Use Not comparable

Inkling: 29.6 (#85), Pixtral Large: —

Agentic & Tool Use benchmarks
BenchmarkInklingPixtral Large
APEX-Agents33.8%—
τ²-bench Banking25%—

Reasoning Inkling leads

Inkling: 40.4 (#56), Pixtral Large: 21.7 (#218)

Reasoning benchmarks
BenchmarkInklingPixtral Large
ARC-AGI-236.5%—
SimpleBench50%—
ARC-AGI-179.5%—
CritPt5.4%—
Chess Puzzles21%—
EnigmaEval—0.8%
LMArena Hard Prompts1451—
DTBench87.5%—
LMCA37.6%—
Epoch Capabilities Index148.54—

Math Not comparable

Inkling: 31.3 (#225), Pixtral Large: —

Math benchmarks
BenchmarkInklingPixtral Large
FrontierMath (Tiers 1-3)33.3%—
FrontierMath Tier 44.9%—
OTIS Mock AIME 2024-202588.9%—
ProofBench0%—
LMArena Math1479—

Knowledge Not comparable

Inkling: 55.1 (#49), Pixtral Large: —

Knowledge benchmarks
BenchmarkInklingPixtral Large
GPQA Diamond88.3%—
SimpleQA Verified40.3%—
LMArena Expert1465—

Multimodal Not comparable

Inkling: —, Pixtral Large: 30.6 (#111)

Multimodal benchmarks
BenchmarkInklingPixtral Large
LMArena Vision—1089

Multilingual Not comparable

Inkling: 54.0 (#52), Pixtral Large: —

Multilingual benchmarks
BenchmarkInklingPixtral Large
LMArena Non-English1434—
LMArena Chinese1490—
LMArena French1458—
LMArena German1446—
LMArena Japanese1429—
LMArena Korean1404—
LMArena Russian1429—
LMArena Spanish1448—

Instruction Following Not comparable

Inkling: 75.1 (#71), Pixtral Large: —

Instruction Following benchmarks
BenchmarkInklingPixtral Large
LMArena Instruction Following1426—

Long Context Not comparable

Inkling: 43.8 (#86), Pixtral Large: —

Long Context benchmarks
BenchmarkInklingPixtral Large
LMArena Longer Query1434—

Writing & Preference Inkling leads

Inkling: 65.2 (#51), Pixtral Large: 32.9 (#278)

Writing & Preference benchmarks
BenchmarkInklingPixtral Large
EQ-Bench Creative Writing1611988
LMArena Text1441—
LMArena Creative Writing1387—
EQ-Bench 41226—
LMArena Multi-Turn1436—

Frequently asked questions

Is Inkling better than Pixtral Large?

Inkling is the stronger model overall, scoring 44.1 to 32.2 on the Noometry Index.

Which is cheaper, Inkling or Pixtral Large?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; Pixtral Large lists at $2 and $6.

Which has the bigger context window?

Pixtral Large does, with 128K tokens against 66K.

How many benchmarks do Inkling and Pixtral Large share?

1 benchmark has published results for both models. Inkling has 41 scored results on Noometry and Pixtral Large has 3.

Related comparisons

Go deeper