Model comparison

Claude 2.1 vs Gemini 1.5 Pro (May 2024)

Gemini 1.5 Pro (May 2024) is the stronger model overall, scoring 32.1 to 25.2 on the Noometry Index.

Last verified . 7 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Gemini 1.5 Pro (May 2024) Google

32.1

Rank #261 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Claude 2.1 scores higher in 1 category and Gemini 1.5 Pro (May 2024) in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemini 1.5 Pro (May 2024) leads 25.8 to 10.2.
  • The biggest single-benchmark swing is GPQA Diamond: 33% for Claude 2.1 and 57.2% for Gemini 1.5 Pro (May 2024).

Side by side

Claude 2.1 and Gemini 1.5 Pro (May 2024) specifications
Claude 2.1Gemini 1.5 Pro (May 2024)
ProviderAnthropicGoogle
Noometry Index25.232.1
Released2023-11-212024-02-15
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked745

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Pro (May 2024) leads

Claude 2.1: 26.2 (#327), Gemini 1.5 Pro (May 2024): 34.2 (#241)

Coding benchmarks
BenchmarkClaude 2.1Gemini 1.5 Pro (May 2024)
WeirdML7.1%22.2%
BigCodeBench Instruct—43.8%
LMArena Coding—1294
BigCodeBench Complete—57.5%
CadEval—34%
HumanEval+—79.3%
MBPP+—74.6%

Agentic & Tool Use Not comparable

Claude 2.1: —, Gemini 1.5 Pro (May 2024): 17.9 (#145)

Agentic & Tool Use benchmarks
BenchmarkClaude 2.1Gemini 1.5 Pro (May 2024)
TheAgentCompany—3.4%
Cybench—7.5%
BALROG—21%

Reasoning Claude 2.1 leads

Claude 2.1: 21.4 (#221), Gemini 1.5 Pro (May 2024): 12.3 (#338)

Reasoning benchmarks
BenchmarkClaude 2.1Gemini 1.5 Pro (May 2024)
DTBench51%59%
Epoch Capabilities Index119.27131.73
ForecastBench54.258.4
ARC-AGI-2—0.8%
SimpleBench—27.1%
LMArena Hard Prompts—1296
BIG-Bench Hard—89.2%

Math Gemini 1.5 Pro (May 2024) leads

Claude 2.1: 10.2 (#315), Gemini 1.5 Pro (May 2024): 25.8 (#266)

Math benchmarks
BenchmarkClaude 2.1Gemini 1.5 Pro (May 2024)
OTIS Mock AIME 2024-20251.9%23.1%
Omni-MATH—36.4%
LMArena Math—1315
MATH Level 5—70.4%

Knowledge Gemini 1.5 Pro (May 2024) leads

Claude 2.1: 15.4 (#292), Gemini 1.5 Pro (May 2024): 29.4 (#239)

Knowledge benchmarks
BenchmarkClaude 2.1Gemini 1.5 Pro (May 2024)
GPQA Diamond33%57.2%
MMLU73.5%86.9%
Humanity's Last Exam—4.6%
MMLU-Pro—73.7%
Confabulations—13.5%
GPQA (HELM)—53.4%
LMArena Expert—1279

Multimodal Not comparable

Claude 2.1: —, Gemini 1.5 Pro (May 2024): 36.8 (#77)

Multimodal benchmarks
BenchmarkClaude 2.1Gemini 1.5 Pro (May 2024)
LMArena Vision—1161
Video-MME—75%

Multilingual Not comparable

Claude 2.1: —, Gemini 1.5 Pro (May 2024): 45.3 (#174)

Multilingual benchmarks
BenchmarkClaude 2.1Gemini 1.5 Pro (May 2024)
LMArena Non-English—1312
LMArena Chinese—1331
LMArena French—1302
LMArena German—1286
LMArena Japanese—1292
LMArena Korean—1298
LMArena Russian—1320
LMArena Spanish—1311

Instruction Following Not comparable

Claude 2.1: —, Gemini 1.5 Pro (May 2024): 68.6 (#185)

Instruction Following benchmarks
BenchmarkClaude 2.1Gemini 1.5 Pro (May 2024)
IFEval—83.7%
LMArena Instruction Following—1297

Long Context Not comparable

Claude 2.1: —, Gemini 1.5 Pro (May 2024): 39.8 (#169)

Long Context benchmarks
BenchmarkClaude 2.1Gemini 1.5 Pro (May 2024)
LMArena Longer Query—1308

Writing & Preference Not comparable

Claude 2.1: —, Gemini 1.5 Pro (May 2024): 52.4 (#172)

Writing & Preference benchmarks
BenchmarkClaude 2.1Gemini 1.5 Pro (May 2024)
LMArena Text—1319
LMArena Creative Writing—1333
WildBench—81.3%
LMArena Multi-Turn—1296

Frequently asked questions

Is Claude 2.1 better than Gemini 1.5 Pro (May 2024)?

Gemini 1.5 Pro (May 2024) is the stronger model overall, scoring 32.1 to 25.2 on the Noometry Index.

Is Claude 2.1 or Gemini 1.5 Pro (May 2024) better for coding?

Gemini 1.5 Pro (May 2024) scores higher on coding benchmarks: 34.2 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Gemini 1.5 Pro (May 2024) share?

7 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Gemini 1.5 Pro (May 2024) has 45.

Related comparisons

Go deeper