Model comparison

Gemma 2 27B vs o1-mini

o1-mini is the stronger model overall, scoring 34.0 to 29.4 on the Noometry Index.

Last verified . 29 shared benchmarks.

Gemma 2 27B Google

29.4

Rank #312 Confirmed

o1-mini OpenAI

34.0

Rank #235 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Gemma 2 27B scores higher in 1 category and o1-mini in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where o1-mini leads 35.4 to 10.7.
  • The biggest single-benchmark swing is MATH Level 5: 27.9% for Gemma 2 27B and 89.2% for o1-mini.
  • Gemma 2 27B has downloadable open weights; the other is API-only.

Side by side

Gemma 2 27B and o1-mini specifications
Gemma 2 27Bo1-mini
ProviderGoogleOpenAI
Noometry Index29.434.0
Released2024-06-242024-09-12
WeightsOpenProprietary
Context window8K—
Max output2K—
Input $ / M tokens$0.65—
Output $ / M tokens$0.65—
Results tracked3439

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1-mini leads

Gemma 2 27B: 34.1 (#246), o1-mini: 35.5 (#224)

Coding benchmarks
BenchmarkGemma 2 27Bo1-mini
LiveBench Coding36%48%
LMArena Coding12111362
Aider Polyglot—32.9%
WeirdML—36.3%
BigCodeBench Instruct42.8%—
BigCodeBench Complete52.5%—
HumanEval+—89%
MBPP+—78.8%

Agentic & Tool Use Not comparable

Gemma 2 27B: —, o1-mini: 24.6 (#118)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 27Bo1-mini
Cybench—10%

Reasoning Gemma 2 27B leads

Gemma 2 27B: 15.3 (#315), o1-mini: 8.8 (#346)

Reasoning benchmarks
BenchmarkGemma 2 27Bo1-mini
LiveBench Reasoning28.1%72.3%
LMArena Hard Prompts11981333
LiveBench Data Analysis47.9%57.9%
Epoch Capabilities Index122.08135.82
LiveBench38.2%57.8%
ARC-AGI-2—0.8%
SimpleBench—18.1%
ARC-AGI-1—14%
DTBench48%—
LMCA7.1%—

Math o1-mini leads

Gemma 2 27B: 10.7 (#311), o1-mini: 35.4 (#186)

Math benchmarks
BenchmarkGemma 2 27Bo1-mini
OTIS Mock AIME 2024-20251.4%46.9%
LiveBench Math26.5%62%
LMArena Math12121358
MATH Level 527.9%89.2%
FrontierMath (Feb 2025 set)—1.7%

Knowledge o1-mini leads

Gemma 2 27B: 19.0 (#280), o1-mini: 34.9 (#192)

Knowledge benchmarks
BenchmarkGemma 2 27Bo1-mini
GPQA Diamond36.5%62.4%
Confabulations27.1%18.6%
LMArena Expert11721316
MMLU75.7%—

Multilingual o1-mini leads

Gemma 2 27B: 38.6 (#226), o1-mini: 43.6 (#182)

Multilingual benchmarks
BenchmarkGemma 2 27Bo1-mini
LMArena Non-English12171289
LMArena Chinese12211314
LMArena French12471293
LMArena German12091278
LMArena Japanese11751245
LMArena Korean11741223
LMArena Russian12341283
LMArena Spanish12281303

Instruction Following o1-mini leads

Gemma 2 27B: 60.5 (#249), o1-mini: 66.7 (#206)

Instruction Following benchmarks
BenchmarkGemma 2 27Bo1-mini
LiveBench Instruction Following58.1%65.4%
LMArena Instruction Following12061304

Long Context o1-mini leads

Gemma 2 27B: 37.3 (#218), o1-mini: 40.1 (#161)

Long Context benchmarks
BenchmarkGemma 2 27Bo1-mini
LMArena Longer Query12311320

Writing & Preference o1-mini leads

Gemma 2 27B: 44.2 (#225), o1-mini: 48.4 (#202)

Writing & Preference benchmarks
BenchmarkGemma 2 27Bo1-mini
LMArena Text12311317
LMArena Creative Writing12411244
LMArena Multi-Turn12241314
LiveBench Language32.6%40.9%
Short-Story Creative Writing—64.9%

Frequently asked questions

Is Gemma 2 27B better than o1-mini?

o1-mini is the stronger model overall, scoring 34.0 to 29.4 on the Noometry Index.

Is Gemma 2 27B or o1-mini better for coding?

o1-mini scores higher on coding benchmarks: 35.5 versus 34.1 in the Noometry coding category.

How many benchmarks do Gemma 2 27B and o1-mini share?

29 benchmarks have published results for both models. Gemma 2 27B has 34 scored results on Noometry and o1-mini has 39.

Related comparisons

Go deeper