Model comparison

Gemma 3 27B vs GPT-4.5

GPT-4.5 is the stronger model overall, scoring 37.2 to 30.8 on the Noometry Index.

Last verified . 33 shared benchmarks.

Gemma 3 27B Google

30.8

Rank #284 Confirmed

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Gemma 3 27B scores higher in 1 category and GPT-4.5 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GPT-4.5 leads 42.2 to 22.5.
  • The biggest single-benchmark swing is Aider Polyglot: 4.9% for Gemma 3 27B and 44.9% for GPT-4.5.
  • Gemma 3 27B has downloadable open weights; the other is API-only.

Side by side

Gemma 3 27B and GPT-4.5 specifications
Gemma 3 27BGPT-4.5
ProviderGoogleOpenAI
Noometry Index30.837.2
Released2025-03-112025-02-27
WeightsOpenProprietary
Context window131K—
Max output8K—
Input $ / M tokens$0.08—
Output $ / M tokens$0.16—
Results tracked4342

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4.5 leads

Gemma 3 27B: 22.5 (#334), GPT-4.5: 42.2 (#109)

Coding benchmarks
BenchmarkGemma 3 27BGPT-4.5
Aider Polyglot4.9%44.9%
LiveBench Coding39.9%75.2%
LMArena Coding13221396
SciCode21.2%—
WeirdML—39.4%

Agentic & Tool Use GPT-4.5 leads

Gemma 3 27B: 25.1 (#110), GPT-4.5: 27.9 (#97)

Agentic & Tool Use benchmarks
BenchmarkGemma 3 27BGPT-4.5
Berkeley Function Calling Leaderboard29.5%—
Cybench—17.5%

Reasoning Gemma 3 27B leads

Gemma 3 27B: 16.7 (#301), GPT-4.5: 13.9 (#330)

Reasoning benchmarks
BenchmarkGemma 3 27BGPT-4.5
LiveBench Reasoning43.8%71.1%
LMArena Hard Prompts13401403
LiveBench Data Analysis51.5%64.3%
Epoch Capabilities Index130.04136.74
LiveBench50%69%
ARC-AGI-2—0.8%
SimpleBench—34.5%
Kagi LLM Benchmark40.4%—
ARC-AGI-1—10.3%
CritPt0%—
Chess Puzzles0%—
EnigmaEval—3.2%
DTBench52.5%—
LMCA12.3%—
ForecastBench—61.7

Math GPT-4.5 leads

Gemma 3 27B: 25.9 (#265), GPT-4.5: 32.6 (#211)

Math benchmarks
BenchmarkGemma 3 27BGPT-4.5
OTIS Mock AIME 2024-202522.5%37.8%
LiveBench Math55.4%69.3%
LMArena Math13121412
MATH Level 574%78.6%

Knowledge GPT-4.5 leads

Gemma 3 27B: 25.5 (#261), GPT-4.5: 32.5 (#211)

Knowledge benchmarks
BenchmarkGemma 3 27BGPT-4.5
GPQA Diamond47.7%68.7%
Confabulations40.3%13.6%
LMArena Expert13041394
Humanity's Last Exam—5.4%
Vectara Hallucination Rate7.4%—

Multimodal GPT-4.5 leads

Gemma 3 27B: 32.6 (#100), GPT-4.5: 37.6 (#71)

Multimodal benchmarks
BenchmarkGemma 3 27BGPT-4.5
LMArena Vision11641195
GeoBench52%—
VPCT—45%

Multilingual GPT-4.5 leads

Gemma 3 27B: 46.9 (#155), GPT-4.5: 52.5 (#83)

Multilingual benchmarks
BenchmarkGemma 3 27BGPT-4.5
LMArena Non-English13341413
LMArena Chinese13461421
LMArena French13681418
LMArena German13621457
LMArena Japanese12871416
LMArena Korean13081392
LMArena Russian13491419
LMArena Spanish1349—

Instruction Following GPT-4.5 leads

Gemma 3 27B: 70.6 (#160), GPT-4.5: 72.6 (#134)

Instruction Following benchmarks
BenchmarkGemma 3 27BGPT-4.5
LiveBench Instruction Following74.9%72.3%
LMArena Instruction Following13211404

Long Context GPT-4.5 leads

Gemma 3 27B: 27.6 (#293), GPT-4.5: 40.4 (#155)

Long Context benchmarks
BenchmarkGemma 3 27BGPT-4.5
Fiction.LiveBench33.3%63.9%
LMArena Longer Query13331406

Writing & Preference GPT-4.5 leads

Gemma 3 27B: 52.5 (#168), GPT-4.5: 56.9 (#134)

Writing & Preference benchmarks
BenchmarkGemma 3 27BGPT-4.5
LMArena Text13581417
LMArena Creative Writing13461394
Short-Story Creative Writing79.9%75.6%
EQ-Bench Creative Writing12661258
LMArena Multi-Turn13451444
LiveBench Language34.6%61.5%

Frequently asked questions

Is Gemma 3 27B better than GPT-4.5?

GPT-4.5 is the stronger model overall, scoring 37.2 to 30.8 on the Noometry Index.

Is Gemma 3 27B or GPT-4.5 better for coding?

GPT-4.5 scores higher on coding benchmarks: 42.2 versus 22.5 in the Noometry coding category.

How many benchmarks do Gemma 3 27B and GPT-4.5 share?

33 benchmarks have published results for both models. Gemma 3 27B has 43 scored results on Noometry and GPT-4.5 has 42.

Related comparisons

Go deeper