Model comparison

Claude Opus 4 vs Gemini 1.5 Pro (May 2024)

Claude Opus 4 is the stronger model overall, scoring 43.1 to 32.1 on the Noometry Index.

Last verified . 35 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

Gemini 1.5 Pro (May 2024) Google

32.1

Rank #261 Confirmed

Summary

  • They share 35 benchmarks with published results for both. Claude Opus 4 scores higher in 8 categories and Gemini 1.5 Pro (May 2024) in 2 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Claude Opus 4 leads 34.8 to 17.9.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 64.4% for Claude Opus 4 and 23.1% for Gemini 1.5 Pro (May 2024).

Side by side

Claude Opus 4 and Gemini 1.5 Pro (May 2024) specifications
Claude Opus 4Gemini 1.5 Pro (May 2024)
ProviderAnthropicGoogle
Noometry Index43.132.1
Released2025-05-222024-02-15
WeightsProprietaryProprietary
Context window200K—
Max output32K—
Input $ / M tokens$15—
Output $ / M tokens$75—
Results tracked5645

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4 leads

Claude Opus 4: 47.2 (#62), Gemini 1.5 Pro (May 2024): 34.2 (#241)

Coding benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Pro (May 2024)
WeirdML43.7%22.2%
LMArena Coding14421294
SWE-bench Verified70.7%—
SWE-bench Verified (bash only)67.6%—
Aider Polyglot72%—
GSO6.9%—
BigCodeBench Instruct—43.8%
BigCodeBench Complete—57.5%
CadEval—34%
AlgoTune1.33—
HumanEval+—79.3%
MBPP+—74.6%

Agentic & Tool Use Claude Opus 4 leads

Claude Opus 4: 34.8 (#42), Gemini 1.5 Pro (May 2024): 17.9 (#145)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Pro (May 2024)
Cybench38%7.5%
TheAgentCompany—3.4%
DeepResearch Bench46.8%—
BALROG—21%
LMArena Search1127—
METR Time Horizons63.9%—

Reasoning Claude Opus 4 leads

Claude Opus 4: 27.3 (#121), Gemini 1.5 Pro (May 2024): 12.3 (#338)

Reasoning benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Pro (May 2024)
ARC-AGI-28.6%0.8%
SimpleBench58.8%27.1%
LMArena Hard Prompts13991296
DTBench81.6%59%
Epoch Capabilities Index142.67131.73
ForecastBench61.158.4
Kagi LLM Benchmark74.3%—
ARC-AGI-135.7%—
CritPt0.3%—
EnigmaEval5.6%—
LMCA37.4%—
BIG-Bench Hard—89.2%

Math Claude Opus 4 leads

Claude Opus 4: 42.0 (#86), Gemini 1.5 Pro (May 2024): 25.8 (#266)

Math benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Pro (May 2024)
OTIS Mock AIME 2024-202564.4%23.1%
Omni-MATH61.6%36.4%
LMArena Math13901315
MATH Level 585%70.4%
FrontierMath (Feb 2025 set)4.5%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Claude Opus 4 leads

Claude Opus 4: 44.0 (#88), Gemini 1.5 Pro (May 2024): 29.4 (#239)

Knowledge benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Pro (May 2024)
GPQA Diamond76.3%57.2%
Humanity's Last Exam10.7%4.6%
MMLU-Pro87.5%73.7%
Confabulations15.9%13.5%
GPQA (HELM)70.8%53.4%
LMArena Expert13861279
Vectara Hallucination Rate12%—
MMLU—86.9%

Multimodal Gemini 1.5 Pro (May 2024) leads

Claude Opus 4: 31.5 (#106), Gemini 1.5 Pro (May 2024): 36.8 (#77)

Multimodal benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Pro (May 2024)
LMArena Vision11921161
Video-MME—75%
GeoBench49%—
VPCT38%—

Multilingual Claude Opus 4 leads

Claude Opus 4: 48.8 (#138), Gemini 1.5 Pro (May 2024): 45.3 (#174)

Multilingual benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Pro (May 2024)
LMArena Non-English13621312
LMArena Chinese13861331
LMArena French13721302
LMArena German13911286
LMArena Japanese13311292
LMArena Korean13211298
LMArena Russian13921320
LMArena Spanish13891311

Instruction Following Claude Opus 4 leads

Claude Opus 4: 77.1 (#28), Gemini 1.5 Pro (May 2024): 68.6 (#185)

Instruction Following benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Pro (May 2024)
IFEval91.8%83.7%
LMArena Instruction Following14061297

Long Context Too close to call

Claude Opus 4: 39.6 (#172), Gemini 1.5 Pro (May 2024): 39.8 (#169)

Long Context benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Pro (May 2024)
LMArena Longer Query14221308
Fiction.LiveBench61.1%—

Writing & Preference Claude Opus 4 leads

Claude Opus 4: 61.2 (#89), Gemini 1.5 Pro (May 2024): 52.4 (#172)

Writing & Preference benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Pro (May 2024)
LMArena Text13771319
LMArena Creative Writing13871333
WildBench85.2%81.3%
LMArena Multi-Turn13961296
Short-Story Creative Writing83.6%—
EQ-Bench Creative Writing1580—

Frequently asked questions

Is Claude Opus 4 better than Gemini 1.5 Pro (May 2024)?

Claude Opus 4 is the stronger model overall, scoring 43.1 to 32.1 on the Noometry Index.

Is Claude Opus 4 or Gemini 1.5 Pro (May 2024) better for coding?

Claude Opus 4 scores higher on coding benchmarks: 47.2 versus 34.2 in the Noometry coding category.

How many benchmarks do Claude Opus 4 and Gemini 1.5 Pro (May 2024) share?

35 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and Gemini 1.5 Pro (May 2024) has 45.

Related comparisons

Go deeper