Model comparison

Gemma 3 27B vs GPT-4o

Gemma 3 27B is the stronger model overall, scoring 30.8 to 28.6 on the Noometry Index.

Last verified . 39 shared benchmarks.

Gemma 3 27B Google

30.8

Rank #284 Confirmed

GPT-4o OpenAI

28.6

Rank #324 Confirmed

Summary

  • They share 39 benchmarks with published results for both. Gemma 3 27B scores higher in 5 categories and GPT-4o in 5 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 3 27B leads 25.9 to 10.6.
  • The biggest single-benchmark swing is Aider Polyglot: 4.9% for Gemma 3 27B and 45.3% for GPT-4o.
  • Gemma 3 27B is cheaper at $0.08 / $0.16 per million input/output tokens, against $2.50 / $10 for GPT-4o.
  • Gemma 3 27B accepts more context: 131K tokens versus 128K.
  • Gemma 3 27B has downloadable open weights; the other is API-only.

Side by side

Gemma 3 27B and GPT-4o specifications
Gemma 3 27BGPT-4o
ProviderGoogleOpenAI
Noometry Index30.828.6
Released2025-03-112024-05-13
WeightsOpenProprietary
Context window131K128K
Max output8K16K
Input $ / M tokens$0.08$2.50
Output $ / M tokens$0.16$10
Results tracked4372

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4o leads

Gemma 3 27B: 22.5 (#334), GPT-4o: 24.8 (#328)

Coding benchmarks
BenchmarkGemma 3 27BGPT-4o
Aider Polyglot4.9%45.3%
LiveBench Coding39.9%51.4%
LMArena Coding13221297
SWE-bench Verified—31%
SWE-bench Verified (bash only)—21.6%
SciCode21.2%—
GSO—0%
WeirdML—25.1%
BigCodeBench Instruct—51.1%
BigCodeBench Complete—61.1%
CadEval—26%
HumanEval+—87.2%
MBPP+—72.2%

Agentic & Tool Use Gemma 3 27B leads

Gemma 3 27B: 25.1 (#110), GPT-4o: 21.0 (#141)

Agentic & Tool Use benchmarks
BenchmarkGemma 3 27BGPT-4o
Berkeley Function Calling Leaderboard29.5%—
GDPval—9.9%
TheAgentCompany—8.6%
Cybench—12.5%
BALROG—32.3%
LMArena Search—1006
METR Time Horizons—40.8%

Reasoning Gemma 3 27B leads

Gemma 3 27B: 16.7 (#301), GPT-4o: 9.4 (#343)

Reasoning benchmarks
BenchmarkGemma 3 27BGPT-4o
CritPt0%0%
Chess Puzzles0%13%
LiveBench Reasoning43.8%55.8%
LMArena Hard Prompts13401281
DTBench52.5%64.5%
LiveBench Data Analysis51.5%60.9%
LMCA12.3%16.6%
Epoch Capabilities Index130.04128.97
LiveBench50%55.3%
ARC-AGI-2—0%
SimpleBench—17.8%
Kagi LLM Benchmark40.4%—
ARC-AGI-1—4.5%
EnigmaEval—0.8%
ForecastBench—57.7

Math Gemma 3 27B leads

Gemma 3 27B: 25.9 (#265), GPT-4o: 10.6 (#312)

Math benchmarks
BenchmarkGemma 3 27BGPT-4o
OTIS Mock AIME 2024-202522.5%6.4%
LiveBench Math55.4%49.5%
LMArena Math13121285
MATH Level 574%53.3%
FrontierMath (Tiers 1-3)—0.4%
Omni-MATH—29.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge GPT-4o leads

Gemma 3 27B: 25.5 (#261), GPT-4o: 28.8 (#242)

Knowledge benchmarks
BenchmarkGemma 3 27BGPT-4o
GPQA Diamond47.7%49.2%
Confabulations40.3%15.3%
Vectara Hallucination Rate7.4%9.6%
LMArena Expert13041250
Humanity's Last Exam—2.7%
SimpleQA Verified—26%
MMLU-Pro—71.3%
GPQA (HELM)—52%
MMLU—88.1%

Multimodal GPT-4o leads

Gemma 3 27B: 32.6 (#100), GPT-4o: 34.5 (#91)

Multimodal benchmarks
BenchmarkGemma 3 27BGPT-4o
LMArena Vision11641137
GeoBench52%71%
Video-MME—71.9%
VPCT—40%
ScienceQA—88.5%

Multilingual Gemma 3 27B leads

Gemma 3 27B: 46.9 (#155), GPT-4o: 43.2 (#186)

Multilingual benchmarks
BenchmarkGemma 3 27BGPT-4o
LMArena Non-English13341283
LMArena Chinese13461277
LMArena French13681304
LMArena German13621282
LMArena Japanese12871257
LMArena Korean13081234
LMArena Russian13491286
LMArena Spanish13491292

Instruction Following Gemma 3 27B leads

Gemma 3 27B: 70.6 (#160), GPT-4o: 66.6 (#207)

Instruction Following benchmarks
BenchmarkGemma 3 27BGPT-4o
LiveBench Instruction Following74.9%68.6%
LMArena Instruction Following13211278
IFEval—81.7%

Long Context GPT-4o leads

Gemma 3 27B: 27.6 (#293), GPT-4o: 39.4 (#179)

Long Context benchmarks
BenchmarkGemma 3 27BGPT-4o
Fiction.LiveBench33.3%66.7%
LMArena Longer Query13331289

Writing & Preference Too close to call

Gemma 3 27B: 52.5 (#168), GPT-4o: 52.6 (#166)

Writing & Preference benchmarks
BenchmarkGemma 3 27BGPT-4o
LMArena Text13581300
LMArena Creative Writing13461292
Short-Story Creative Writing79.9%81.8%
LMArena Multi-Turn13451302
LiveBench Language34.6%47.6%
EQ-Bench Creative Writing1266—
WildBench—82.8%

Frequently asked questions

Is Gemma 3 27B better than GPT-4o?

Gemma 3 27B is the stronger model overall, scoring 30.8 to 28.6 on the Noometry Index.

Which is cheaper, Gemma 3 27B or GPT-4o?

Gemma 3 27B is cheaper. It lists at $0.08 per million input tokens and $0.16 per million output tokens; GPT-4o lists at $2.50 and $10.

Is Gemma 3 27B or GPT-4o better for coding?

GPT-4o scores higher on coding benchmarks: 24.8 versus 22.5 in the Noometry coding category.

Which has the bigger context window?

Gemma 3 27B does, with 131K tokens against 128K.

How many benchmarks do Gemma 3 27B and GPT-4o share?

39 benchmarks have published results for both models. Gemma 3 27B has 43 scored results on Noometry and GPT-4o has 72.

Related comparisons

Go deeper