Model comparison

Codellama 70b Instruct vs Inkling

Inkling is the stronger model overall, scoring 44.1 to 33.7 on the Noometry Index.

Last verified . 4 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Codellama 70b Instruct scores higher in 1 category and Inkling in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Inkling leads 65.2 to 33.4.

Side by side

Codellama 70b Instruct and Inkling specifications
Codellama 70b InstructInkling
ProviderMetaThinking Machines Lab
Noometry Index33.744.1
Released—2026-07-15
WeightsOpenOpen
Context window—66K
Max output—66K
Input $ / M tokens—$1.87
Output $ / M tokens—$4.68
Results tracked741

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 70b Instruct leads

Codellama 70b Instruct: 37.6 (#193), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkCodellama 70b InstructInkling
FrontierCode—14%
LMArena WebDev—1413
FrontierSWE—4.1%
SciCode—47%
WeirdML—32.3%
BigCodeBench Instruct40.7%—
LMArena Coding—1464
BigCodeBench Complete49.6%—
ALE-Bench—946
HumanEval+65.9%—

Agentic & Tool Use Not comparable

Codellama 70b Instruct: —, Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkCodellama 70b InstructInkling
APEX-Agents—33.8%
τ²-bench Banking—25%

Reasoning Inkling leads

Codellama 70b Instruct: 20.1 (#242), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkCodellama 70b InstructInkling
LMArena Hard Prompts10521451
ARC-AGI-2—36.5%
SimpleBench—50%
ARC-AGI-1—79.5%
CritPt—5.4%
Chess Puzzles—21%
DTBench—87.5%
LMCA—37.6%
Epoch Capabilities Index—148.54

Math Not comparable

Codellama 70b Instruct: —, Inkling: 31.3 (#225)

Math benchmarks
BenchmarkCodellama 70b InstructInkling
FrontierMath (Tiers 1-3)—33.3%
FrontierMath Tier 4—4.9%
OTIS Mock AIME 2024-2025—88.9%
ProofBench—0%
LMArena Math—1479

Knowledge Not comparable

Codellama 70b Instruct: —, Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkCodellama 70b InstructInkling
GPQA Diamond—88.3%
SimpleQA Verified—40.3%
LMArena Expert—1465

Multilingual Inkling leads

Codellama 70b Instruct: 24.8 (#288), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkCodellama 70b InstructInkling
LMArena Non-English9921434
LMArena Chinese—1490
LMArena French—1458
LMArena German—1446
LMArena Japanese—1429
LMArena Korean—1404
LMArena Russian—1429
LMArena Spanish—1448

Instruction Following Inkling leads

Codellama 70b Instruct: 51.9 (#293), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructInkling
LMArena Instruction Following10241426

Long Context Not comparable

Codellama 70b Instruct: —, Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkCodellama 70b InstructInkling
LMArena Longer Query—1434

Writing & Preference Inkling leads

Codellama 70b Instruct: 33.4 (#277), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructInkling
LMArena Text10571441
LMArena Creative Writing—1387
EQ-Bench Creative Writing—1611
EQ-Bench 4—1226
LMArena Multi-Turn—1436

Frequently asked questions

Is Codellama 70b Instruct better than Inkling?

Inkling is the stronger model overall, scoring 44.1 to 33.7 on the Noometry Index.

Is Codellama 70b Instruct or Inkling better for coding?

Codellama 70b Instruct scores higher on coding benchmarks: 37.6 versus 34.5 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Inkling share?

4 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper