Model comparison

Gemma 1.1 7b IT vs o1

o1 is the stronger model overall, scoring 40.9 to 31.3 on the Noometry Index.

Last verified . 19 shared benchmarks.

Gemma 1.1 7b IT Google

31.3

Rank #277 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Gemma 1.1 7b IT scores higher in 0 categories and o1 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where o1 leads 55.6 to 30.4.
  • Gemma 1.1 7b IT has downloadable open weights; the other is API-only.

Side by side

Gemma 1.1 7b IT and o1 specifications
Gemma 1.1 7b ITo1
ProviderGoogleOpenAI
Noometry Index31.340.9
Released—2024-09-12
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked1952

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Gemma 1.1 7b IT: 31.5 (#284), o1: 46.1 (#70)

Coding benchmarks
BenchmarkGemma 1.1 7b ITo1
LMArena Coding10841367
HumanEval+35.4%89%
MBPP+45%80.2%
Aider Polyglot—61.7%
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%

Agentic & Tool Use Not comparable

Gemma 1.1 7b IT: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkGemma 1.1 7b ITo1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Gemma 1.1 7b IT: 20.5 (#238), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkGemma 1.1 7b ITo1
LMArena Hard Prompts10711371
SimpleBench—41.7%
ARC-AGI-1—30.7%
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
Epoch Capabilities Index—141.91
LiveBench—75.7%

Math o1 leads

Gemma 1.1 7b IT: 32.0 (#220), o1: 36.1 (#175)

Math benchmarks
BenchmarkGemma 1.1 7b ITo1
LMArena Math11071388
FrontierMath (Tiers 1-3)—14.7%
OTIS Mock AIME 2024-2025—73.3%
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

Gemma 1.1 7b IT: 28.3 (#247), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkGemma 1.1 7b ITo1
LMArena Expert10391361
GPQA Diamond—76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%

Multimodal Not comparable

Gemma 1.1 7b IT: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkGemma 1.1 7b ITo1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual o1 leads

Gemma 1.1 7b IT: 28.1 (#273), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkGemma 1.1 7b ITo1
LMArena Non-English10521358
LMArena Chinese10611394
LMArena French10651344
LMArena German10541337
LMArena Japanese9711346
LMArena Korean9881396
LMArena Russian10461356
LMArena Spanish10491345

Instruction Following o1 leads

Gemma 1.1 7b IT: 54.0 (#283), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkGemma 1.1 7b ITo1
LMArena Instruction Following10571367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Gemma 1.1 7b IT: 32.1 (#272), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkGemma 1.1 7b ITo1
LMArena Longer Query10561378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

Gemma 1.1 7b IT: 30.4 (#288), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkGemma 1.1 7b ITo1
LMArena Text10941366
LMArena Creative Writing10601348
LMArena Multi-Turn10401369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is Gemma 1.1 7b IT better than o1?

o1 is the stronger model overall, scoring 40.9 to 31.3 on the Noometry Index.

Is Gemma 1.1 7b IT or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 31.5 in the Noometry coding category.

How many benchmarks do Gemma 1.1 7b IT and o1 share?

19 benchmarks have published results for both models. Gemma 1.1 7b IT has 19 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper