Model comparison

Gemma 1.1 2b IT vs o1

o1 is the stronger model overall, scoring 40.9 to 29.3 on the Noometry Index.

Last verified . 16 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 0 categories and o1 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where o1 leads 55.6 to 25.1.
  • Gemma 1.1 2b IT has downloadable open weights; the other is API-only.

Side by side

Gemma 1.1 2b IT and o1 specifications
Gemma 1.1 2b ITo1
ProviderGoogleOpenAI
Noometry Index29.340.9
Released—2024-09-12
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked1652

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Gemma 1.1 2b IT: 30.1 (#299), o1: 46.1 (#70)

Coding benchmarks
BenchmarkGemma 1.1 2b ITo1
LMArena Coding10341367
HumanEval+17.7%89%
MBPP+23.3%80.2%
Aider Polyglot—61.7%
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%

Agentic & Tool Use Not comparable

Gemma 1.1 2b IT: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkGemma 1.1 2b ITo1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Gemma 1.1 2b IT: 19.1 (#270), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITo1
LMArena Hard Prompts10051371
SimpleBench—41.7%
ARC-AGI-1—30.7%
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
Epoch Capabilities Index—141.91
LiveBench—75.7%

Math o1 leads

Gemma 1.1 2b IT: 30.8 (#232), o1: 36.1 (#175)

Math benchmarks
BenchmarkGemma 1.1 2b ITo1
LMArena Math10471388
FrontierMath (Tiers 1-3)—14.7%
OTIS Mock AIME 2024-2025—73.3%
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

Gemma 1.1 2b IT: 26.5 (#258), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITo1
LMArena Expert9701361
GPQA Diamond—76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%

Multimodal Not comparable

Gemma 1.1 2b IT: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkGemma 1.1 2b ITo1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual o1 leads

Gemma 1.1 2b IT: 24.6 (#289), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITo1
LMArena Non-English9881358
LMArena Chinese10121394
LMArena German9441337
LMArena Korean8991396
LMArena Russian9901356
LMArena French—1344
LMArena Japanese—1346
LMArena Spanish—1345

Instruction Following o1 leads

Gemma 1.1 2b IT: 49.9 (#299), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITo1
LMArena Instruction Following9921367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Gemma 1.1 2b IT: 30.6 (#286), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITo1
LMArena Longer Query10031378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

Gemma 1.1 2b IT: 25.1 (#306), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITo1
LMArena Text10221366
LMArena Creative Writing9981348
LMArena Multi-Turn9591369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is Gemma 1.1 2b IT better than o1?

o1 is the stronger model overall, scoring 40.9 to 29.3 on the Noometry Index.

Is Gemma 1.1 2b IT or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 30.1 in the Noometry coding category.

How many benchmarks do Gemma 1.1 2b IT and o1 share?

16 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper