Model comparison

GPT-4 Turbo vs Inkling

Inkling is the stronger model overall, scoring 44.1 to 30.5 on the Noometry Index.

Last verified . 26 shared benchmarks.

GPT-4 Turbo OpenAI

30.5

Rank #292 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 26 benchmarks with published results for both. GPT-4 Turbo scores higher in 0 categories and Inkling in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 24.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 6.7% for GPT-4 Turbo and 88.9% for Inkling.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $10 / $30 for GPT-4 Turbo.
  • GPT-4 Turbo accepts more context: 128K tokens versus 66K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

GPT-4 Turbo and Inkling specifications
GPT-4 TurboInkling
ProviderOpenAIThinking Machines Lab
Noometry Index30.544.1
Released2023-11-062026-07-15
WeightsProprietaryOpen
Context window128K66K
Max output4K66K
Input $ / M tokens$10$1.87
Output $ / M tokens$30$4.68
Results tracked3641

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-4 Turbo: 33.8 (#249), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGPT-4 TurboInkling
WeirdML18%32.3%
LMArena Coding12681464
FrontierCode—14%
LMArena WebDev—1413
FrontierSWE—4.1%
SciCode—47%
BigCodeBench Instruct48.2%—
BigCodeBench Complete58.2%—
ALE-Bench—946
HumanEval+86.6%—
MBPP+73.3%—

Agentic & Tool Use Not comparable

GPT-4 Turbo: —, Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGPT-4 TurboInkling
APEX-Agents—33.8%
τ²-bench Banking—25%
METR Time Horizons36.7%—

Reasoning Inkling leads

GPT-4 Turbo: 15.3 (#317), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGPT-4 TurboInkling
SimpleBench25.1%50%
Chess Puzzles6%21%
LMArena Hard Prompts12511451
DTBench61.6%87.5%
LMCA9.8%37.6%
Epoch Capabilities Index127.25148.54
ARC-AGI-2—36.5%
ARC-AGI-1—79.5%
CritPt—5.4%
ForecastBench59.4—

Math Inkling leads

GPT-4 Turbo: 9.0 (#322), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGPT-4 TurboInkling
FrontierMath (Tiers 1-3)0.7%33.3%
OTIS Mock AIME 2024-20256.7%88.9%
LMArena Math12721479
FrontierMath Tier 4—4.9%
ProofBench—0%
MATH Level 546.7%—

Knowledge Inkling leads

GPT-4 Turbo: 24.3 (#268), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGPT-4 TurboInkling
GPQA Diamond46.6%88.3%
LMArena Expert12231465
SimpleQA Verified—40.3%
Confabulations28.4%—
MMLU81.3%—

Multimodal Not comparable

GPT-4 Turbo: 30.6 (#110), Inkling: —

Multimodal benchmarks
BenchmarkGPT-4 TurboInkling
LMArena Vision1090—

Multilingual Inkling leads

GPT-4 Turbo: 40.5 (#216), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGPT-4 TurboInkling
LMArena Non-English12451434
LMArena Chinese12421490
LMArena French12761458
LMArena German12591446
LMArena Japanese11941429
LMArena Korean11871404
LMArena Russian12591429
LMArena Spanish12601448

Instruction Following Inkling leads

GPT-4 Turbo: 65.8 (#216), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGPT-4 TurboInkling
LMArena Instruction Following12491426

Long Context Inkling leads

GPT-4 Turbo: 38.0 (#206), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGPT-4 TurboInkling
LMArena Longer Query12541434

Writing & Preference Inkling leads

GPT-4 Turbo: 47.7 (#206), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGPT-4 TurboInkling
LMArena Text12721441
LMArena Creative Writing12691387
LMArena Multi-Turn12671436
EQ-Bench Creative Writing—1611
EQ-Bench 4—1226

Frequently asked questions

Is GPT-4 Turbo better than Inkling?

Inkling is the stronger model overall, scoring 44.1 to 30.5 on the Noometry Index.

Which is cheaper, GPT-4 Turbo or Inkling?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; GPT-4 Turbo lists at $10 and $30.

Is GPT-4 Turbo or Inkling better for coding?

They score almost the same on coding (33.8 vs 34.5); test both on your own repository before choosing.

Which has the bigger context window?

GPT-4 Turbo does, with 128K tokens against 66K.

How many benchmarks do GPT-4 Turbo and Inkling share?

26 benchmarks have published results for both models. GPT-4 Turbo has 36 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper