Model comparison

Gemini 1.0 Pro vs GPT-3.5-turbo

Gemini 1.0 Pro is the stronger model overall, scoring 27.3 to 23.2 on the Noometry Index.

Last verified . 24 shared benchmarks.

Gemini 1.0 Pro Google

27.3

Rank #332 Confirmed

GPT-3.5-turbo OpenAI

23.2

Rank #350 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Gemini 1.0 Pro scores higher in 7 categories and GPT-3.5-turbo in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemini 1.0 Pro leads 36.0 to 25.3.
  • The biggest single-benchmark swing is GPQA Diamond: 34% for Gemini 1.0 Pro and 28% for GPT-3.5-turbo.

Side by side

Gemini 1.0 Pro and GPT-3.5-turbo specifications
Gemini 1.0 ProGPT-3.5-turbo
ProviderGoogleOpenAI
Noometry Index27.323.2
Released2023-12-132023-03-01
WeightsProprietaryProprietary
Context window—16K
Max output—4K
Input $ / M tokens—$0.50
Output $ / M tokens—$1.50
Results tracked2444

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.0 Pro leads

Gemini 1.0 Pro: 32.2 (#275), GPT-3.5-turbo: 23.9 (#331)

Coding benchmarks
BenchmarkGemini 1.0 ProGPT-3.5-turbo
LMArena Coding11081136
HumanEval+55.5%70.7%
MBPP+61.4%69.7%
WeirdML—3.5%
BigCodeBench Instruct—39.1%
BigCodeBench Complete—50.6%

Agentic & Tool Use Not comparable

Gemini 1.0 Pro: —, GPT-3.5-turbo: —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.0 ProGPT-3.5-turbo
METR Time Horizons—21.5%

Reasoning Gemini 1.0 Pro leads

Gemini 1.0 Pro: 17.1 (#296), GPT-3.5-turbo: 13.8 (#332)

Reasoning benchmarks
BenchmarkGemini 1.0 ProGPT-3.5-turbo
LMArena Hard Prompts11091108
DTBench45.9%48.5%
Epoch Capabilities Index117.04118.55
Chess Puzzles—0%
Mystery Game Puzzles—3%
LMCA—9.7%
Adversarial NLI—58.1%
BIG-Bench Hard—61.6%
CommonsenseQA 2.0—57%
ForecastBench—50.4
WinoGrande—81.6%

Math Gemini 1.0 Pro leads

Gemini 1.0 Pro: 9.3 (#321), GPT-3.5-turbo: 6.3 (#327)

Math benchmarks
BenchmarkGemini 1.0 ProGPT-3.5-turbo
OTIS Mock AIME 2024-20251.1%2.2%
LMArena Math11321142
MATH Level 511.2%15.9%
FrontierMath (Tiers 1-3)—0%
GSM8K—57.8%

Knowledge Gemini 1.0 Pro leads

Gemini 1.0 Pro: 15.6 (#291), GPT-3.5-turbo: 10.0 (#303)

Knowledge benchmarks
BenchmarkGemini 1.0 ProGPT-3.5-turbo
GPQA Diamond34%28%
LMArena Expert10591070
MMLU70%71.4%
ARC (AI2) Challenge—87.4%
BoolQ—87%
OpenBookQA—86%
TriviaQA—85.8%

Multilingual Gemini 1.0 Pro leads

Gemini 1.0 Pro: 33.4 (#252), GPT-3.5-turbo: 31.5 (#258)

Multilingual benchmarks
BenchmarkGemini 1.0 ProGPT-3.5-turbo
LMArena Non-English11381108
LMArena Chinese11241075
LMArena French11451118
LMArena German11251090
LMArena Japanese10231043
LMArena Russian11861123
LMArena Spanish11191121
LMArena Korean—1019

Instruction Following Too close to call

Gemini 1.0 Pro: 57.6 (#267), GPT-3.5-turbo: 57.9 (#262)

Instruction Following benchmarks
BenchmarkGemini 1.0 ProGPT-3.5-turbo
LMArena Instruction Following11141119

Long Context Too close to call

Gemini 1.0 Pro: 34.3 (#249), GPT-3.5-turbo: 34.0 (#254)

Long Context benchmarks
BenchmarkGemini 1.0 ProGPT-3.5-turbo
LMArena Longer Query11321121

Writing & Preference Gemini 1.0 Pro leads

Gemini 1.0 Pro: 36.0 (#264), GPT-3.5-turbo: 25.3 (#305)

Writing & Preference benchmarks
BenchmarkGemini 1.0 ProGPT-3.5-turbo
LMArena Text11491125
LMArena Creative Writing11311092
LMArena Multi-Turn11391117
EQ-Bench Creative Writing—451

Frequently asked questions

Is Gemini 1.0 Pro better than GPT-3.5-turbo?

Gemini 1.0 Pro is the stronger model overall, scoring 27.3 to 23.2 on the Noometry Index.

Is Gemini 1.0 Pro or GPT-3.5-turbo better for coding?

Gemini 1.0 Pro scores higher on coding benchmarks: 32.2 versus 23.9 in the Noometry coding category.

How many benchmarks do Gemini 1.0 Pro and GPT-3.5-turbo share?

24 benchmarks have published results for both models. Gemini 1.0 Pro has 24 scored results on Noometry and GPT-3.5-turbo has 44.

Related comparisons

Go deeper