Model comparison

Gemma 7B vs GPT-4

Gemma 7B and GPT-4 score almost the same on the Noometry Index (30.0 vs 29.1), so choose on price, context window or the category you care about most.

Last verified . 21 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Gemma 7B scores higher in 3 categories and GPT-4 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 7B leads 31.2 to 10.8.
  • Gemma 7B has downloadable open weights; the other is API-only.

Side by side

Gemma 7B and GPT-4 specifications
Gemma 7BGPT-4
ProviderGoogleOpenAI
Noometry Index30.029.1
Released2024-02-212023-03-14
WeightsOpenProprietary
Context window—8K
Max output—8K
Input $ / M tokens—$30
Output $ / M tokens—$60
Results tracked2738

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4 leads

Gemma 7B: 30.5 (#294), GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkGemma 7BGPT-4
LMArena Coding10481254
HumanEval+28.7%79.3%
WeirdML—12.4%
BigCodeBench Instruct—46%
BigCodeBench Complete—57.2%
MBPP+43.4%—

Agentic & Tool Use Not comparable

Gemma 7B: —, GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkGemma 7BGPT-4
METR Time Horizons—36.1%

Reasoning Gemma 7B leads

Gemma 7B: 19.9 (#249), GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkGemma 7BGPT-4
LMArena Hard Prompts10421241
BIG-Bench Hard55.1%75.1%
Epoch Capabilities Index111.99125.89
HellaSwag82.2%95.3%
WinoGrande79%87.5%
Chess Puzzles—4%
Mystery Game Puzzles—12%
DTBench—62.7%
LMCA—17.1%
Adversarial NLI48.7%—
ForecastBench—57.8
PIQA81.2%—

Math Gemma 7B leads

Gemma 7B: 31.2 (#228), GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkGemma 7BGPT-4
LMArena Math10661269
GSM8K46.4%92%
OTIS Mock AIME 2024-2025—1.1%
MATH Level 5—23%

Knowledge Gemma 7B leads

Gemma 7B: 27.3 (#252), GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkGemma 7BGPT-4
LMArena Expert10011211
MMLU66.1%86.4%
TriviaQA72.3%84.8%
GPQA Diamond—35.7%
ARC (AI2) Challenge78.3%—
BoolQ83.2%—
OpenBookQA78.6%—

Multilingual GPT-4 leads

Gemma 7B: 25.1 (#287), GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkGemma 7BGPT-4
LMArena Non-English9991246
LMArena Chinese10351242
LMArena French10251283
LMArena Russian9931251
LMArena German—1251
LMArena Japanese—1209
LMArena Korean—1184
LMArena Spanish—1261

Instruction Following GPT-4 leads

Gemma 7B: 51.5 (#295), GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkGemma 7BGPT-4
LMArena Instruction Following10171241

Long Context GPT-4 leads

Gemma 7B: 31.1 (#282), GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkGemma 7BGPT-4
LMArena Longer Query10221244

Writing & Preference GPT-4 leads

Gemma 7B: 27.1 (#302), GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkGemma 7BGPT-4
LMArena Text10561263
LMArena Creative Writing10241244
LMArena Multi-Turn9631257
EQ-Bench Creative Writing—752

Frequently asked questions

Is Gemma 7B better than GPT-4?

Gemma 7B and GPT-4 score almost the same on the Noometry Index (30.0 vs 29.1), so choose on price, context window or the category you care about most.

Is Gemma 7B or GPT-4 better for coding?

GPT-4 scores higher on coding benchmarks: 31.6 versus 30.5 in the Noometry coding category.

How many benchmarks do Gemma 7B and GPT-4 share?

21 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper