Model comparison

Gemma 1.1 2b IT vs GPT-4

Gemma 1.1 2b IT and GPT-4 score almost the same on the Noometry Index (29.3 vs 29.1), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 3 categories and GPT-4 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 1.1 2b IT leads 30.8 to 10.8.
  • Gemma 1.1 2b IT has downloadable open weights; the other is API-only.

Side by side

Gemma 1.1 2b IT and GPT-4 specifications
Gemma 1.1 2b ITGPT-4
ProviderGoogleOpenAI
Noometry Index29.329.1
Released—2023-03-14
WeightsOpenProprietary
Context window—8K
Max output—8K
Input $ / M tokens—$30
Output $ / M tokens—$60
Results tracked1638

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4 leads

Gemma 1.1 2b IT: 30.1 (#299), GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkGemma 1.1 2b ITGPT-4
LMArena Coding10341254
HumanEval+17.7%79.3%
WeirdML—12.4%
BigCodeBench Instruct—46%
BigCodeBench Complete—57.2%
MBPP+23.3%—

Agentic & Tool Use Not comparable

Gemma 1.1 2b IT: —, GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkGemma 1.1 2b ITGPT-4
METR Time Horizons—36.1%

Reasoning Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 19.1 (#270), GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITGPT-4
LMArena Hard Prompts10051241
Chess Puzzles—4%
Mystery Game Puzzles—12%
DTBench—62.7%
LMCA—17.1%
BIG-Bench Hard—75.1%
Epoch Capabilities Index—125.89
ForecastBench—57.8
HellaSwag—95.3%
WinoGrande—87.5%

Math Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 30.8 (#232), GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkGemma 1.1 2b ITGPT-4
LMArena Math10471269
OTIS Mock AIME 2024-2025—1.1%
MATH Level 5—23%
GSM8K—92%

Knowledge Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 26.5 (#258), GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITGPT-4
LMArena Expert9701211
GPQA Diamond—35.7%
MMLU—86.4%
TriviaQA—84.8%

Multilingual GPT-4 leads

Gemma 1.1 2b IT: 24.6 (#289), GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITGPT-4
LMArena Non-English9881246
LMArena Chinese10121242
LMArena German9441251
LMArena Korean8991184
LMArena Russian9901251
LMArena French—1283
LMArena Japanese—1209
LMArena Spanish—1261

Instruction Following GPT-4 leads

Gemma 1.1 2b IT: 49.9 (#299), GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITGPT-4
LMArena Instruction Following9921241

Long Context GPT-4 leads

Gemma 1.1 2b IT: 30.6 (#286), GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITGPT-4
LMArena Longer Query10031244

Writing & Preference GPT-4 leads

Gemma 1.1 2b IT: 25.1 (#306), GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITGPT-4
LMArena Text10221263
LMArena Creative Writing9981244
LMArena Multi-Turn9591257
EQ-Bench Creative Writing—752

Frequently asked questions

Is Gemma 1.1 2b IT better than GPT-4?

Gemma 1.1 2b IT and GPT-4 score almost the same on the Noometry Index (29.3 vs 29.1), so choose on price, context window or the category you care about most.

Is Gemma 1.1 2b IT or GPT-4 better for coding?

GPT-4 scores higher on coding benchmarks: 31.6 versus 30.1 in the Noometry coding category.

How many benchmarks do Gemma 1.1 2b IT and GPT-4 share?

15 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper