Model comparison

Gemma 3 27B vs GPT-3.5-turbo

Gemma 3 27B is the stronger model overall, scoring 30.8 to 23.2 on the Noometry Index.

Last verified . 25 shared benchmarks.

Gemma 3 27B Google

30.8

Rank #284 Confirmed

GPT-3.5-turbo OpenAI

23.2

Rank #350 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Gemma 3 27B scores higher in 6 categories and GPT-3.5-turbo in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemma 3 27B leads 52.5 to 25.3.
  • The biggest single-benchmark swing is MATH Level 5: 74% for Gemma 3 27B and 15.9% for GPT-3.5-turbo.
  • Gemma 3 27B is cheaper at $0.08 / $0.16 per million input/output tokens, against $0.50 / $1.50 for GPT-3.5-turbo.
  • Gemma 3 27B accepts more context: 131K tokens versus 16K.
  • Gemma 3 27B has downloadable open weights; the other is API-only.

Side by side

Gemma 3 27B and GPT-3.5-turbo specifications
Gemma 3 27BGPT-3.5-turbo
ProviderGoogleOpenAI
Noometry Index30.823.2
Released2025-03-112023-03-01
WeightsOpenProprietary
Context window131K16K
Max output8K4K
Input $ / M tokens$0.08$0.50
Output $ / M tokens$0.16$1.50
Results tracked4344

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-3.5-turbo leads

Gemma 3 27B: 22.5 (#334), GPT-3.5-turbo: 23.9 (#331)

Coding benchmarks
BenchmarkGemma 3 27BGPT-3.5-turbo
LMArena Coding13221136
Aider Polyglot4.9%—
SciCode21.2%—
WeirdML—3.5%
BigCodeBench Instruct—39.1%
LiveBench Coding39.9%—
BigCodeBench Complete—50.6%
HumanEval+—70.7%
MBPP+—69.7%

Agentic & Tool Use Not comparable

Gemma 3 27B: 25.1 (#110), GPT-3.5-turbo: —

Agentic & Tool Use benchmarks
BenchmarkGemma 3 27BGPT-3.5-turbo
Berkeley Function Calling Leaderboard29.5%—
METR Time Horizons—21.5%

Reasoning Gemma 3 27B leads

Gemma 3 27B: 16.7 (#301), GPT-3.5-turbo: 13.8 (#332)

Reasoning benchmarks
BenchmarkGemma 3 27BGPT-3.5-turbo
Chess Puzzles0%0%
LMArena Hard Prompts13401108
DTBench52.5%48.5%
LMCA12.3%9.7%
Epoch Capabilities Index130.04118.55
Kagi LLM Benchmark40.4%—
CritPt0%—
LiveBench Reasoning43.8%—
Mystery Game Puzzles—3%
LiveBench Data Analysis51.5%—
Adversarial NLI—58.1%
BIG-Bench Hard—61.6%
CommonsenseQA 2.0—57%
ForecastBench—50.4
LiveBench50%—
WinoGrande—81.6%

Math Gemma 3 27B leads

Gemma 3 27B: 25.9 (#265), GPT-3.5-turbo: 6.3 (#327)

Math benchmarks
BenchmarkGemma 3 27BGPT-3.5-turbo
OTIS Mock AIME 2024-202522.5%2.2%
LMArena Math13121142
MATH Level 574%15.9%
FrontierMath (Tiers 1-3)—0%
LiveBench Math55.4%—
GSM8K—57.8%

Knowledge Gemma 3 27B leads

Gemma 3 27B: 25.5 (#261), GPT-3.5-turbo: 10.0 (#303)

Knowledge benchmarks
BenchmarkGemma 3 27BGPT-3.5-turbo
GPQA Diamond47.7%28%
LMArena Expert13041070
Confabulations40.3%—
Vectara Hallucination Rate7.4%—
ARC (AI2) Challenge—87.4%
BoolQ—87%
MMLU—71.4%
OpenBookQA—86%
TriviaQA—85.8%

Multimodal Not comparable

Gemma 3 27B: 32.6 (#100), GPT-3.5-turbo: —

Multimodal benchmarks
BenchmarkGemma 3 27BGPT-3.5-turbo
LMArena Vision1164—
GeoBench52%—

Multilingual Gemma 3 27B leads

Gemma 3 27B: 46.9 (#155), GPT-3.5-turbo: 31.5 (#258)

Multilingual benchmarks
BenchmarkGemma 3 27BGPT-3.5-turbo
LMArena Non-English13341108
LMArena Chinese13461075
LMArena French13681118
LMArena German13621090
LMArena Japanese12871043
LMArena Korean13081019
LMArena Russian13491123
LMArena Spanish13491121

Instruction Following Gemma 3 27B leads

Gemma 3 27B: 70.6 (#160), GPT-3.5-turbo: 57.9 (#262)

Instruction Following benchmarks
BenchmarkGemma 3 27BGPT-3.5-turbo
LMArena Instruction Following13211119
LiveBench Instruction Following74.9%—

Long Context GPT-3.5-turbo leads

Gemma 3 27B: 27.6 (#293), GPT-3.5-turbo: 34.0 (#254)

Long Context benchmarks
BenchmarkGemma 3 27BGPT-3.5-turbo
LMArena Longer Query13331121
Fiction.LiveBench33.3%—

Writing & Preference Gemma 3 27B leads

Gemma 3 27B: 52.5 (#168), GPT-3.5-turbo: 25.3 (#305)

Writing & Preference benchmarks
BenchmarkGemma 3 27BGPT-3.5-turbo
LMArena Text13581125
LMArena Creative Writing13461092
EQ-Bench Creative Writing1266451
LMArena Multi-Turn13451117
Short-Story Creative Writing79.9%—
LiveBench Language34.6%—

Frequently asked questions

Is Gemma 3 27B better than GPT-3.5-turbo?

Gemma 3 27B is the stronger model overall, scoring 30.8 to 23.2 on the Noometry Index.

Which is cheaper, Gemma 3 27B or GPT-3.5-turbo?

Gemma 3 27B is cheaper. It lists at $0.08 per million input tokens and $0.16 per million output tokens; GPT-3.5-turbo lists at $0.50 and $1.50.

Is Gemma 3 27B or GPT-3.5-turbo better for coding?

GPT-3.5-turbo scores higher on coding benchmarks: 23.9 versus 22.5 in the Noometry coding category.

Which has the bigger context window?

Gemma 3 27B does, with 131K tokens against 16K.

How many benchmarks do Gemma 3 27B and GPT-3.5-turbo share?

25 benchmarks have published results for both models. Gemma 3 27B has 43 scored results on Noometry and GPT-3.5-turbo has 44.

Related comparisons

Go deeper