Model comparison

Claude 3 Opus vs Gemini 1.5 Pro (May 2024)

Gemini 1.5 Pro (May 2024) is the stronger model overall, scoring 32.1 to 29.5 on the Noometry Index.

Last verified . 33 shared benchmarks.

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

Gemini 1.5 Pro (May 2024) Google

32.1

Rank #261 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Claude 3 Opus scores higher in 2 categories and Gemini 1.5 Pro (May 2024) in 8 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemini 1.5 Pro (May 2024) leads 25.8 to 14.8.
  • The biggest single-benchmark swing is MATH Level 5: 37.5% for Claude 3 Opus and 70.4% for Gemini 1.5 Pro (May 2024).

Side by side

Claude 3 Opus and Gemini 1.5 Pro (May 2024) specifications
Claude 3 OpusGemini 1.5 Pro (May 2024)
ProviderAnthropicGoogle
Noometry Index29.532.1
Released2024-02-292024-02-15
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4645

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Pro (May 2024) leads

Claude 3 Opus: 32.9 (#267), Gemini 1.5 Pro (May 2024): 34.2 (#241)

Coding benchmarks
BenchmarkClaude 3 OpusGemini 1.5 Pro (May 2024)
WeirdML19.2%22.2%
BigCodeBench Instruct45.5%43.8%
LMArena Coding12641294
BigCodeBench Complete57.4%57.5%
HumanEval+77.4%79.3%
MBPP+73.3%74.6%
LiveBench Coding38.6%—
CadEval—34%

Agentic & Tool Use Claude 3 Opus leads

Claude 3 Opus: 24.6 (#116), Gemini 1.5 Pro (May 2024): 17.9 (#145)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 OpusGemini 1.5 Pro (May 2024)
Cybench10%7.5%
TheAgentCompany—3.4%
BALROG—21%
METR Time Horizons29.5%—

Reasoning Claude 3 Opus leads

Claude 3 Opus: 14.6 (#324), Gemini 1.5 Pro (May 2024): 12.3 (#338)

Reasoning benchmarks
BenchmarkClaude 3 OpusGemini 1.5 Pro (May 2024)
SimpleBench23.5%27.1%
LMArena Hard Prompts12451296
DTBench61.6%59%
Epoch Capabilities Index126.91131.73
ForecastBench58.458.4
ARC-AGI-2—0.8%
Chess Puzzles5%—
EnigmaEval0.8%—
LiveBench Reasoning40.6%—
LiveBench Data Analysis57.9%—
LMCA17%—
BIG-Bench Hard—89.2%
LiveBench49.2%—
WinoGrande88.5%—

Math Gemini 1.5 Pro (May 2024) leads

Claude 3 Opus: 14.8 (#299), Gemini 1.5 Pro (May 2024): 25.8 (#266)

Math benchmarks
BenchmarkClaude 3 OpusGemini 1.5 Pro (May 2024)
OTIS Mock AIME 2024-20254.7%23.1%
LMArena Math12731315
MATH Level 537.5%70.4%
Omni-MATH—36.4%
LiveBench Math43.6%—

Knowledge Gemini 1.5 Pro (May 2024) leads

Claude 3 Opus: 24.5 (#267), Gemini 1.5 Pro (May 2024): 29.4 (#239)

Knowledge benchmarks
BenchmarkClaude 3 OpusGemini 1.5 Pro (May 2024)
GPQA Diamond47.2%57.2%
Confabulations22.7%13.5%
LMArena Expert12231279
MMLU84.6%86.9%
Humanity's Last Exam—4.6%
SimpleQA Verified12.6%—
MMLU-Pro—73.7%
GPQA (HELM)—53.4%

Multimodal Gemini 1.5 Pro (May 2024) leads

Claude 3 Opus: 27.1 (#116), Gemini 1.5 Pro (May 2024): 36.8 (#77)

Multimodal benchmarks
BenchmarkClaude 3 OpusGemini 1.5 Pro (May 2024)
LMArena Vision10231161
Video-MME—75%

Multilingual Gemini 1.5 Pro (May 2024) leads

Claude 3 Opus: 41.4 (#207), Gemini 1.5 Pro (May 2024): 45.3 (#174)

Multilingual benchmarks
BenchmarkClaude 3 OpusGemini 1.5 Pro (May 2024)
LMArena Non-English12581312
LMArena Chinese12481331
LMArena French12751302
LMArena German12581286
LMArena Japanese12041292
LMArena Korean11871298
LMArena Russian12801320
LMArena Spanish12461311

Instruction Following Gemini 1.5 Pro (May 2024) leads

Claude 3 Opus: 64.1 (#228), Gemini 1.5 Pro (May 2024): 68.6 (#185)

Instruction Following benchmarks
BenchmarkClaude 3 OpusGemini 1.5 Pro (May 2024)
LMArena Instruction Following12481297
LiveBench Instruction Following63.9%—
IFEval—83.7%

Long Context Gemini 1.5 Pro (May 2024) leads

Claude 3 Opus: 38.2 (#202), Gemini 1.5 Pro (May 2024): 39.8 (#169)

Long Context benchmarks
BenchmarkClaude 3 OpusGemini 1.5 Pro (May 2024)
LMArena Longer Query12591308

Writing & Preference Gemini 1.5 Pro (May 2024) leads

Claude 3 Opus: 47.2 (#213), Gemini 1.5 Pro (May 2024): 52.4 (#172)

Writing & Preference benchmarks
BenchmarkClaude 3 OpusGemini 1.5 Pro (May 2024)
LMArena Text12621319
LMArena Creative Writing12351333
LMArena Multi-Turn12751296
WildBench—81.3%
LiveBench Language50.4%—

Frequently asked questions

Is Claude 3 Opus better than Gemini 1.5 Pro (May 2024)?

Gemini 1.5 Pro (May 2024) is the stronger model overall, scoring 32.1 to 29.5 on the Noometry Index.

Is Claude 3 Opus or Gemini 1.5 Pro (May 2024) better for coding?

Gemini 1.5 Pro (May 2024) scores higher on coding benchmarks: 34.2 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3 Opus and Gemini 1.5 Pro (May 2024) share?

33 benchmarks have published results for both models. Claude 3 Opus has 46 scored results on Noometry and Gemini 1.5 Pro (May 2024) has 45.

Related comparisons

Go deeper