Model comparison

Gemma 3n E4b IT vs Inkling

Inkling is the stronger model overall, scoring 44.1 to 37.3 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 3n E4b IT scores higher in 2 categories and Inkling in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 34.2.

Side by side

Gemma 3n E4b IT and Inkling specifications
Gemma 3n E4b ITInkling
ProviderGoogleThinking Machines Lab
Noometry Index37.344.1
Released—2026-07-15
WeightsOpenOpen
Context window—66K
Max output—66K
Input $ / M tokens—$1.87
Output $ / M tokens—$4.68
Results tracked1841

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 3n E4b IT leads

Gemma 3n E4b IT: 37.0 (#198), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGemma 3n E4b ITInkling
LMArena Coding12681464
FrontierCode—14%
LMArena WebDev—1413
FrontierSWE—4.1%
SciCode—47%
WeirdML—32.3%
ALE-Bench—946

Agentic & Tool Use Not comparable

Gemma 3n E4b IT: —, Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGemma 3n E4b ITInkling
APEX-Agents—33.8%
τ²-bench Banking—25%

Reasoning Inkling leads

Gemma 3n E4b IT: 19.9 (#247), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGemma 3n E4b ITInkling
LMArena Hard Prompts12841451
ARC-AGI-2—36.5%
SimpleBench—50%
Kagi LLM Benchmark31.5%—
ARC-AGI-1—79.5%
CritPt—5.4%
Chess Puzzles—21%
DTBench—87.5%
LMCA—37.6%
Epoch Capabilities Index—148.54

Math Gemma 3n E4b IT leads

Gemma 3n E4b IT: 35.1 (#188), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGemma 3n E4b ITInkling
LMArena Math12511479
FrontierMath (Tiers 1-3)—33.3%
FrontierMath Tier 4—4.9%
OTIS Mock AIME 2024-2025—88.9%
ProofBench—0%

Knowledge Inkling leads

Gemma 3n E4b IT: 34.2 (#198), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGemma 3n E4b ITInkling
LMArena Expert12461465
GPQA Diamond—88.3%
SimpleQA Verified—40.3%

Multilingual Inkling leads

Gemma 3n E4b IT: 43.4 (#183), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGemma 3n E4b ITInkling
LMArena Non-English12851434
LMArena Chinese13091490
LMArena French13301458
LMArena German13111446
LMArena Japanese12721429
LMArena Korean12591404
LMArena Russian12881429
LMArena Spanish13051448

Instruction Following Inkling leads

Gemma 3n E4b IT: 66.1 (#210), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGemma 3n E4b ITInkling
LMArena Instruction Following12551426

Long Context Inkling leads

Gemma 3n E4b IT: 38.7 (#191), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGemma 3n E4b ITInkling
LMArena Longer Query12761434

Writing & Preference Inkling leads

Gemma 3n E4b IT: 50.1 (#186), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGemma 3n E4b ITInkling
LMArena Text13061441
LMArena Creative Writing12871387
LMArena Multi-Turn12761436
EQ-Bench Creative Writing—1611
EQ-Bench 4—1226

Frequently asked questions

Is Gemma 3n E4b IT better than Inkling?

Inkling is the stronger model overall, scoring 44.1 to 37.3 on the Noometry Index.

Is Gemma 3n E4b IT or Inkling better for coding?

Gemma 3n E4b IT scores higher on coding benchmarks: 37.0 versus 34.5 in the Noometry coding category.

How many benchmarks do Gemma 3n E4b IT and Inkling share?

17 benchmarks have published results for both models. Gemma 3n E4b IT has 18 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper