Model comparison

gpt-oss-20b vs Inkling-Small

Inkling-Small is the stronger model overall, scoring 46.5 to 32.5 on the Noometry Index. gpt-oss-20b costs 18× less per token, which makes it the better buy when Inkling-Small's lead doesn't matter for your workload.

Last verified . 23 shared benchmarks.

gpt-oss-20b OpenAI

32.5

Rank #255 Confirmed

Inkling-Small Thinking Machines Lab

46.5

Rank #63 Confirmed

Summary

  • They share 23 benchmarks with published results for both. gpt-oss-20b scores higher in 0 categories and Inkling-Small in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Inkling-Small leads 59.6 to 35.5.
  • The biggest single-benchmark swing is GPQA Diamond: 60.8% for gpt-oss-20b and 88.5% for Inkling-Small.
  • gpt-oss-20b is cheaper at $0.018 / $0.09 per million input/output tokens, against $0.45 / $1.20 for Inkling-Small.
  • Inkling-Small accepts more context: 524K tokens versus 131K.

Side by side

gpt-oss-20b and Inkling-Small specifications
gpt-oss-20bInkling-Small
ProviderOpenAIThinking Machines Lab
Noometry Index32.546.5
Released2025-08-052026-07-15
WeightsOpenOpen
Context window131K524K
Max output16K1.05M
Input $ / M tokens$0.018$0.45
Output $ / M tokens$0.09$1.20
Results tracked3433

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling-Small leads

gpt-oss-20b: 37.6 (#192), Inkling-Small: 43.6 (#85)

Coding benchmarks
Benchmarkgpt-oss-20bInkling-Small
SciCode34.4%48.7%
LMArena Coding13061451
LMArena WebDev—1409
WeirdML40.9%—
ALE-Bench566.05—

Agentic & Tool Use Not comparable

gpt-oss-20b: 9.3 (#154), Inkling-Small: —

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-20bInkling-Small
Terminal-Bench3.4%—

Reasoning Inkling-Small leads

gpt-oss-20b: 19.3 (#261), Inkling-Small: 38.6 (#63)

Reasoning benchmarks
Benchmarkgpt-oss-20bInkling-Small
CritPt1.4%8.3%
Chess Puzzles4%18%
LMArena Hard Prompts12741423
Epoch Capabilities Index137.82150.15
ARC-AGI-2—40.1%
Kagi LLM Benchmark53.2%—
ARC-AGI-1—84%
Mystery Game Puzzles—6%
DTBench68%—
LMCA14.5%—

Math Inkling-Small leads

gpt-oss-20b: 39.4 (#103), Inkling-Small: 45.1 (#77)

Math benchmarks
Benchmarkgpt-oss-20bInkling-Small
OTIS Mock AIME 2024-202565.3%90%
LMArena Math13171459
FrontierMath (Tiers 1-3)—46.3%
FrontierMath Tier 4—17.1%
ProofBench—6%
Omni-MATH56.5%—

Knowledge Inkling-Small leads

gpt-oss-20b: 34.6 (#195), Inkling-Small: 48.2 (#77)

Knowledge benchmarks
Benchmarkgpt-oss-20bInkling-Small
GPQA Diamond60.8%88.5%
LMArena Expert12581442
SimpleQA Verified—19.1%
MMLU-Pro74%—
GPQA (HELM)59.4%—

Multimodal Not comparable

gpt-oss-20b: —, Inkling-Small: 39.1 (#62)

Multimodal benchmarks
Benchmarkgpt-oss-20bInkling-Small
LMArena Vision—1235

Multilingual Inkling-Small leads

gpt-oss-20b: 42.2 (#197), Inkling-Small: 51.7 (#104)

Multilingual benchmarks
Benchmarkgpt-oss-20bInkling-Small
LMArena Non-English12681402
LMArena Chinese13141465
LMArena German12551405
LMArena Japanese12441405
LMArena Korean12361363
LMArena Russian12781391
LMArena Spanish12671428
LMArena French—1436

Instruction Following Inkling-Small leads

gpt-oss-20b: 61.8 (#240), Inkling-Small: 73.8 (#114)

Instruction Following benchmarks
Benchmarkgpt-oss-20bInkling-Small
LMArena Instruction Following12361399
IFEval73.2%—

Long Context Inkling-Small leads

gpt-oss-20b: 37.9 (#209), Inkling-Small: 42.7 (#118)

Long Context benchmarks
Benchmarkgpt-oss-20bInkling-Small
LMArena Longer Query12501401

Writing & Preference Inkling-Small leads

gpt-oss-20b: 35.5 (#265), Inkling-Small: 59.6 (#107)

Writing & Preference benchmarks
Benchmarkgpt-oss-20bInkling-Small
LMArena Text12871414
LMArena Creative Writing12011331
EQ-Bench Creative Writing6661491
LMArena Multi-Turn12681418
WildBench73.7%—

Frequently asked questions

Is gpt-oss-20b better than Inkling-Small?

Inkling-Small is the stronger model overall, scoring 46.5 to 32.5 on the Noometry Index. gpt-oss-20b costs 18× less per token, which makes it the better buy when Inkling-Small's lead doesn't matter for your workload.

Which is cheaper, gpt-oss-20b or Inkling-Small?

gpt-oss-20b is cheaper. It lists at $0.018 per million input tokens and $0.09 per million output tokens; Inkling-Small lists at $0.45 and $1.20.

Is gpt-oss-20b or Inkling-Small better for coding?

Inkling-Small scores higher on coding benchmarks: 43.6 versus 37.6 in the Noometry coding category.

Which has the bigger context window?

Inkling-Small does, with 524K tokens against 131K.

How many benchmarks do gpt-oss-20b and Inkling-Small share?

23 benchmarks have published results for both models. gpt-oss-20b has 34 scored results on Noometry and Inkling-Small has 33.

Related comparisons

Go deeper