Model comparison

Claude Opus 4 vs Gemini 1.5 Flash (May 2024)

Claude Opus 4 is the stronger model overall, scoring 43.1 to 33.2 on the Noometry Index.

Last verified . 32 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Summary

  • They share 32 benchmarks with published results for both. Claude Opus 4 scores higher in 9 categories and Gemini 1.5 Flash (May 2024) in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Opus 4 leads 42.0 to 22.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 64.4% for Claude Opus 4 and 16.3% for Gemini 1.5 Flash (May 2024).

Side by side

Claude Opus 4 and Gemini 1.5 Flash (May 2024) specifications
Claude Opus 4Gemini 1.5 Flash (May 2024)
ProviderAnthropicGoogle
Noometry Index43.133.2
Released2025-05-222024-05-14
WeightsProprietaryProprietary
Context window200K—
Max output32K—
Input $ / M tokens$15—
Output $ / M tokens$75—
Results tracked5642

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4 leads

Claude Opus 4: 47.2 (#62), Gemini 1.5 Flash (May 2024): 34.4 (#236)

Coding benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Flash (May 2024)
WeirdML43.7%24.9%
LMArena Coding14421261
SWE-bench Verified70.7%—
SWE-bench Verified (bash only)67.6%—
Aider Polyglot72%—
GSO6.9%—
BigCodeBench Instruct—43.5%
BigCodeBench Complete—55.1%
AlgoTune1.33—
HumanEval+—75.6%
MBPP+—67.5%

Agentic & Tool Use Claude Opus 4 leads

Claude Opus 4: 34.8 (#42), Gemini 1.5 Flash (May 2024): 26.6 (#102)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Flash (May 2024)
Cybench38%—
DeepResearch Bench46.8%—
BALROG—14.6%
LMArena Search1127—
METR Time Horizons63.9%—

Reasoning Claude Opus 4 leads

Claude Opus 4: 27.3 (#121), Gemini 1.5 Flash (May 2024): 21.7 (#215)

Reasoning benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Flash (May 2024)
LMArena Hard Prompts13991257
DTBench81.6%53.8%
Epoch Capabilities Index142.67129.36
ForecastBench61.153.9
ARC-AGI-28.6%—
SimpleBench58.8%—
Kagi LLM Benchmark74.3%—
ARC-AGI-135.7%—
CritPt0.3%—
EnigmaEval5.6%—
LMCA37.4%—
PIQA—87.5%

Math Claude Opus 4 leads

Claude Opus 4: 42.0 (#86), Gemini 1.5 Flash (May 2024): 22.1 (#281)

Math benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Flash (May 2024)
OTIS Mock AIME 2024-202564.4%16.3%
Omni-MATH61.6%30.4%
LMArena Math13901269
MATH Level 585%61.9%
FrontierMath (Feb 2025 set)4.5%0%
FrontierMath Tier 4 (v1)4.2%—
GSM8K—82.4%

Knowledge Claude Opus 4 leads

Claude Opus 4: 44.0 (#88), Gemini 1.5 Flash (May 2024): 26.2 (#260)

Knowledge benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Flash (May 2024)
GPQA Diamond76.3%47.3%
MMLU-Pro87.5%67.8%
GPQA (HELM)70.8%43.7%
LMArena Expert13861233
Humanity's Last Exam10.7%—
Confabulations15.9%—
Vectara Hallucination Rate12%—
BoolQ—85.8%
MMLU—77.9%

Multimodal Gemini 1.5 Flash (May 2024) leads

Claude Opus 4: 31.5 (#106), Gemini 1.5 Flash (May 2024): 36.0 (#81)

Multimodal benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Flash (May 2024)
LMArena Vision11921141
GeoBench49%76%
Video-MME—70.3%
VPCT38%—

Multilingual Claude Opus 4 leads

Claude Opus 4: 48.8 (#138), Gemini 1.5 Flash (May 2024): 42.9 (#189)

Multilingual benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Flash (May 2024)
LMArena Non-English13621278
LMArena Chinese13861295
LMArena French13721258
LMArena German13911262
LMArena Japanese13311252
LMArena Korean13211221
LMArena Russian13921288
LMArena Spanish13891243

Instruction Following Claude Opus 4 leads

Claude Opus 4: 77.1 (#28), Gemini 1.5 Flash (May 2024): 66.8 (#205)

Instruction Following benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Flash (May 2024)
IFEval91.8%83.1%
LMArena Instruction Following14061258

Long Context Too close to call

Claude Opus 4: 39.6 (#172), Gemini 1.5 Flash (May 2024): 39.0 (#187)

Long Context benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Flash (May 2024)
LMArena Longer Query14221284
Fiction.LiveBench61.1%—

Writing & Preference Claude Opus 4 leads

Claude Opus 4: 61.2 (#89), Gemini 1.5 Flash (May 2024): 48.7 (#196)

Writing & Preference benchmarks
BenchmarkClaude Opus 4Gemini 1.5 Flash (May 2024)
LMArena Text13771287
LMArena Creative Writing13871285
WildBench85.2%79.2%
LMArena Multi-Turn13961253
Short-Story Creative Writing83.6%—
EQ-Bench Creative Writing1580—

Frequently asked questions

Is Claude Opus 4 better than Gemini 1.5 Flash (May 2024)?

Claude Opus 4 is the stronger model overall, scoring 43.1 to 33.2 on the Noometry Index.

Is Claude Opus 4 or Gemini 1.5 Flash (May 2024) better for coding?

Claude Opus 4 scores higher on coding benchmarks: 47.2 versus 34.4 in the Noometry coding category.

How many benchmarks do Claude Opus 4 and Gemini 1.5 Flash (May 2024) share?

32 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and Gemini 1.5 Flash (May 2024) has 42.

Related comparisons

Go deeper