Model comparison

Gemma 2 27B vs GPT-5 Nano

GPT-5 Nano is the stronger model overall, scoring 33.5 to 29.4 on the Noometry Index.

Last verified . 22 shared benchmarks.

Gemma 2 27B Google

29.4

Rank #312 Confirmed

GPT-5 Nano OpenAI

33.5

Rank #241 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Gemma 2 27B scores higher in 3 categories and GPT-5 Nano in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5 Nano leads 29.4 to 10.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 1.4% for Gemma 2 27B and 81.1% for GPT-5 Nano.
  • GPT-5 Nano is cheaper at $0.05 / $0.40 per million input/output tokens, against $0.65 / $0.65 for Gemma 2 27B.
  • GPT-5 Nano accepts more context: 400K tokens versus 8K.
  • Gemma 2 27B has downloadable open weights; the other is API-only.

Side by side

Gemma 2 27B and GPT-5 Nano specifications
Gemma 2 27BGPT-5 Nano
ProviderGoogleOpenAI
Noometry Index29.433.5
Released2024-06-242025-08-07
WeightsOpenProprietary
Context window8K400K
Max output2K128K
Input $ / M tokens$0.65$0.05
Output $ / M tokens$0.65$0.40
Results tracked3449

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 2 27B: 34.1 (#246), GPT-5 Nano: 33.6 (#254)

Coding benchmarks
BenchmarkGemma 2 27BGPT-5 Nano
LMArena Coding12111351
SWE-bench Verified (bash only)—34.8%
WeirdML—38.1%
BigCodeBench Instruct42.8%—
LiveBench Coding36%—
BigCodeBench Complete52.5%—
ALE-Bench—718.67

Agentic & Tool Use Not comparable

Gemma 2 27B: —, GPT-5 Nano: 25.8 (#106)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 27BGPT-5 Nano
Terminal-Bench—21.8%
Berkeley Function Calling Leaderboard—51.5%

Reasoning Too close to call

Gemma 2 27B: 15.3 (#315), GPT-5 Nano: 16.3 (#306)

Reasoning benchmarks
BenchmarkGemma 2 27BGPT-5 Nano
LMArena Hard Prompts11981328
DTBench48%62.7%
LMCA7.1%7.9%
Epoch Capabilities Index122.08139.38
ARC-AGI-2—2.6%
Kagi LLM Benchmark—62.2%
ARC-AGI-1—20.7%
Chess Puzzles—27%
LiveBench Reasoning28.1%—
Mystery Game Puzzles—9%
LiveBench Data Analysis47.9%—
ForecastBench—59.1
LiveBench38.2%—

Math GPT-5 Nano leads

Gemma 2 27B: 10.7 (#311), GPT-5 Nano: 29.4 (#241)

Math benchmarks
BenchmarkGemma 2 27BGPT-5 Nano
OTIS Mock AIME 2024-20251.4%81.1%
LMArena Math12121317
MATH Level 527.9%95.2%
FrontierMath (Tiers 1-3)—20%
FrontierMath Tier 4—2.4%
ProofBench—12%
Omni-MATH—54.6%
LiveBench Math26.5%—
FrontierMath (Feb 2025 set)—8.3%
FrontierMath Tier 4 (v1)—2.1%

Knowledge GPT-5 Nano leads

Gemma 2 27B: 19.0 (#280), GPT-5 Nano: 35.9 (#178)

Knowledge benchmarks
BenchmarkGemma 2 27BGPT-5 Nano
GPQA Diamond36.5%69.4%
LMArena Expert11721321
SimpleQA Verified—11.7%
MMLU-Pro—77.8%
Confabulations27.1%—
Vectara Hallucination Rate—10.5%
GPQA (HELM)—67.9%
MMLU75.7%—

Multimodal Not comparable

Gemma 2 27B: —, GPT-5 Nano: 31.3 (#108)

Multimodal benchmarks
BenchmarkGemma 2 27BGPT-5 Nano
LMArena Vision—1159
VPCT—37.2%

Multilingual GPT-5 Nano leads

Gemma 2 27B: 38.6 (#226), GPT-5 Nano: 45.3 (#172)

Multilingual benchmarks
BenchmarkGemma 2 27BGPT-5 Nano
LMArena Non-English12171313
LMArena Chinese12211356
LMArena German12091327
LMArena Japanese11751226
LMArena Korean11741269
LMArena Russian12341296
LMArena Spanish12281360
LMArena French1247—

Instruction Following GPT-5 Nano leads

Gemma 2 27B: 60.5 (#249), GPT-5 Nano: 75.0 (#79)

Instruction Following benchmarks
BenchmarkGemma 2 27BGPT-5 Nano
LMArena Instruction Following12061306
LiveBench Instruction Following58.1%—
IFEval—93.2%

Long Context Gemma 2 27B leads

Gemma 2 27B: 37.3 (#218), GPT-5 Nano: 31.3 (#281)

Long Context benchmarks
BenchmarkGemma 2 27BGPT-5 Nano
LMArena Longer Query12311312
Fiction.LiveBench—44.4%

Writing & Preference Gemma 2 27B leads

Gemma 2 27B: 44.2 (#225), GPT-5 Nano: 39.1 (#249)

Writing & Preference benchmarks
BenchmarkGemma 2 27BGPT-5 Nano
LMArena Text12311320
LMArena Creative Writing12411249
LMArena Multi-Turn12241311
EQ-Bench Creative Writing—705
WildBench—80.6%
LiveBench Language32.6%—

Frequently asked questions

Is Gemma 2 27B better than GPT-5 Nano?

GPT-5 Nano is the stronger model overall, scoring 33.5 to 29.4 on the Noometry Index.

Which is cheaper, Gemma 2 27B or GPT-5 Nano?

GPT-5 Nano is cheaper. It lists at $0.05 per million input tokens and $0.40 per million output tokens; Gemma 2 27B lists at $0.65 and $0.65.

Is Gemma 2 27B or GPT-5 Nano better for coding?

They score almost the same on coding (34.1 vs 33.6); test both on your own repository before choosing.

Which has the bigger context window?

GPT-5 Nano does, with 400K tokens against 8K.

How many benchmarks do Gemma 2 27B and GPT-5 Nano share?

22 benchmarks have published results for both models. Gemma 2 27B has 34 scored results on Noometry and GPT-5 Nano has 49.

Related comparisons

Go deeper