Model comparison

Claude Opus 4.6 vs Gemini 1.5 Flash (May 2024)

Claude Opus 4.6 is the stronger model overall, scoring 58.2 to 33.2 on the Noometry Index.

Last verified . 25 shared benchmarks.

Claude Opus 4.6 Anthropic

58.2

Rank #20 Confirmed

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude Opus 4.6 scores higher in 10 categories and Gemini 1.5 Flash (May 2024) in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Opus 4.6 leads 63.0 to 22.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 94.4% for Claude Opus 4.6 and 16.3% for Gemini 1.5 Flash (May 2024).

Side by side

Claude Opus 4.6 and Gemini 1.5 Flash (May 2024) specifications
Claude Opus 4.6Gemini 1.5 Flash (May 2024)
ProviderAnthropicGoogle
Noometry Index58.233.2
Released2026-02-042024-05-14
WeightsProprietaryProprietary
Context window1M—
Max output128K—
Input $ / M tokens$5—
Output $ / M tokens$25—
Results tracked6842

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.6 leads

Claude Opus 4.6: 57.2 (#20), Gemini 1.5 Flash (May 2024): 34.4 (#236)

Coding benchmarks
BenchmarkClaude Opus 4.6Gemini 1.5 Flash (May 2024)
WeirdML78%24.9%
LMArena Coding15361261
SWE-bench Verified78.7%—
FrontierCode26.6%—
SWE-bench Verified (bash only)75.6%—
LMArena WebDev1547—
SWE-bench Multilingual72%—
GSO41.2%—
BigCodeBench Instruct—43.5%
BigCodeBench Complete—55.1%
ALE-Bench996.5—
AlgoTune1.47—
HumanEval+—75.6%
MBPP+—67.5%

Agentic & Tool Use Claude Opus 4.6 leads

Claude Opus 4.6: 51.1 (#4), Gemini 1.5 Flash (May 2024): 26.6 (#102)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.6Gemini 1.5 Flash (May 2024)
Terminal-Bench79.8%—
APEX-Agents46.3%—
Remote Labor Index4.2%—
τ²-bench Banking27.3%—
Cybench93%—
DeepResearch Bench55.3%—
BALROG—14.6%
GBAEval44.1%—
LMArena Search1253—
METR Time Horizons78.9%—
Vending-Bench 28,018—

Reasoning Claude Opus 4.6 leads

Claude Opus 4.6: 57.8 (#23), Gemini 1.5 Flash (May 2024): 21.7 (#215)

Reasoning benchmarks
BenchmarkClaude Opus 4.6Gemini 1.5 Flash (May 2024)
LMArena Hard Prompts15271257
DTBench91.2%53.8%
Epoch Capabilities Index155.24129.36
ForecastBench6053.9
ARC-AGI-269.2%—
SimpleBench67.6%—
Kagi LLM Benchmark83.6%—
NYT Connections (extended)92.1%—
ARC-AGI-194%—
Chess Puzzles17%—
EnigmaEval7.6%—
Thematic Generalization80.6%—
EBR-Bench12.7%—
Mystery Game Puzzles25%—
LMCA55.8%—
PIQA—87.5%

Math Claude Opus 4.6 leads

Claude Opus 4.6: 63.0 (#31), Gemini 1.5 Flash (May 2024): 22.1 (#281)

Math benchmarks
BenchmarkClaude Opus 4.6Gemini 1.5 Flash (May 2024)
OTIS Mock AIME 2024-202594.4%16.3%
LMArena Math15191269
FrontierMath (Feb 2025 set)40.7%0%
FrontierMath (Tiers 1-3)66%—
FrontierMath Tier 426.8%—
MathArena Final-Answer Competitions78.5%—
ProofBench50%—
Omni-MATH—30.4%
MATH Level 5—61.9%
FrontierMath Tier 4 (v1)22.9%—
GSM8K—82.4%

Knowledge Claude Opus 4.6 leads

Claude Opus 4.6: 61.9 (#26), Gemini 1.5 Flash (May 2024): 26.2 (#260)

Knowledge benchmarks
BenchmarkClaude Opus 4.6Gemini 1.5 Flash (May 2024)
GPQA Diamond90.5%47.3%
LMArena Expert15461233
Humanity's Last Exam34.4%—
SimpleQA Verified47%—
MMLU-Pro—67.8%
Vectara Hallucination Rate12.2%—
GPQA (HELM)—43.7%
BoolQ—85.8%
MMLU—77.9%

Multimodal Claude Opus 4.6 leads

Claude Opus 4.6: 37.3 (#74), Gemini 1.5 Flash (May 2024): 36.0 (#81)

Multimodal benchmarks
BenchmarkClaude Opus 4.6Gemini 1.5 Flash (May 2024)
LMArena Vision13161141
Video-MME—70.3%
GeoBench—76%
Furniture Assembly28.3%—
LMArena Document1507—

Multilingual Claude Opus 4.6 leads

Claude Opus 4.6: 57.9 (#6), Gemini 1.5 Flash (May 2024): 42.9 (#189)

Multilingual benchmarks
BenchmarkClaude Opus 4.6Gemini 1.5 Flash (May 2024)
LMArena Non-English14891278
LMArena Chinese15511295
LMArena French15131258
LMArena German15021262
LMArena Japanese14841252
LMArena Korean14641221
LMArena Russian14971288
LMArena Spanish15101243

Instruction Following Claude Opus 4.6 leads

Claude Opus 4.6: 79.5 (#4), Gemini 1.5 Flash (May 2024): 66.8 (#205)

Instruction Following benchmarks
BenchmarkClaude Opus 4.6Gemini 1.5 Flash (May 2024)
LMArena Instruction Following15231258
IFEval—83.1%

Long Context Claude Opus 4.6 leads

Claude Opus 4.6: 48.1 (#13), Gemini 1.5 Flash (May 2024): 39.0 (#187)

Long Context benchmarks
BenchmarkClaude Opus 4.6Gemini 1.5 Flash (May 2024)
LMArena Longer Query15201284
CL-bench20.7%—
CL-bench Life17%—

Writing & Preference Claude Opus 4.6 leads

Claude Opus 4.6: 73.5 (#10), Gemini 1.5 Flash (May 2024): 48.7 (#196)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.6Gemini 1.5 Flash (May 2024)
LMArena Text15031287
LMArena Creative Writing15051285
LMArena Multi-Turn15131253
EQ-Bench Creative Writing1809—
WildBench—79.2%
EQ-Bench 41223—

Frequently asked questions

Is Claude Opus 4.6 better than Gemini 1.5 Flash (May 2024)?

Claude Opus 4.6 is the stronger model overall, scoring 58.2 to 33.2 on the Noometry Index.

Is Claude Opus 4.6 or Gemini 1.5 Flash (May 2024) better for coding?

Claude Opus 4.6 scores higher on coding benchmarks: 57.2 versus 34.4 in the Noometry coding category.

How many benchmarks do Claude Opus 4.6 and Gemini 1.5 Flash (May 2024) share?

25 benchmarks have published results for both models. Claude Opus 4.6 has 68 scored results on Noometry and Gemini 1.5 Flash (May 2024) has 42.

Related comparisons

Go deeper