Model comparison

Claude Opus 4.7 vs Gemini 3 Pro

Claude Opus 4.7 is the stronger model overall, scoring 58.3 to 54.8 on the Noometry Index.

Last verified . 46 shared benchmarks.

Claude Opus 4.7 Anthropic

58.3

Rank #19 Confirmed

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

Summary

  • They share 46 benchmarks with published results for both. Claude Opus 4.7 scores higher in 8 categories and Gemini 3 Pro in 2 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Opus 4.7 leads 66.7 to 49.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 39% for Claude Opus 4.7 and 94.4% for Gemini 3 Pro.

Side by side

Claude Opus 4.7 and Gemini 3 Pro specifications
Claude Opus 4.7Gemini 3 Pro
ProviderAnthropicGoogle
Noometry Index58.354.8
Released2026-04-142025-11-18
WeightsProprietaryProprietary
Context window1M—
Max output128K—
Input $ / M tokens$5—
Output $ / M tokens$25—
Results tracked6667

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.7 leads

Claude Opus 4.7: 59.6 (#13), Gemini 3 Pro: 51.6 (#39)

Coding benchmarks
BenchmarkClaude Opus 4.7Gemini 3 Pro
SWE-bench Verified83.5%72.9%
LMArena WebDev15581440
GSO44.1%18.6%
WeirdML76.4%69.9%
LMArena Coding15181481
ALE-Bench1,3231,177
FrontierCode38.5%—
SWE-bench Verified (bash only)—74.2%
SWE-bench Multilingual—68.7%
SciCode54.5%—
MirrorCode31.1%—
AlgoTune—1.83

Agentic & Tool Use Claude Opus 4.7 leads

Claude Opus 4.7: 47.9 (#10), Gemini 3 Pro: 40.6 (#23)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.7Gemini 3 Pro
Terminal-Bench80.2%69.4%
τ²-bench Banking40.2%18%
LMArena Search12331207
Vending-Bench 210,9375,478
APEX-Agents49.2%—
Berkeley Function Calling Leaderboard—72.5%
OSWorld 2.018.2%—
GDPval—40.3%
Remote Labor Index—1.3%
τ²-bench Airline—80.5%
τ²-bench Retail—75.9%
τ²-bench Telecom—91%
DeepResearch Bench—46.3%
PostTrainBench28.6%—
BALROG—58.1%
ExploitBench26.5%—
GBAEval43.8%—
GDP.pdf21%—
METR Time Horizons—71%

Reasoning Claude Opus 4.7 leads

Claude Opus 4.7: 53.8 (#29), Gemini 3 Pro: 52.5 (#31)

Reasoning benchmarks
BenchmarkClaude Opus 4.7Gemini 3 Pro
ARC-AGI-275.8%31.1%
SimpleBench61.7%76.4%
Kagi LLM Benchmark80.7%80.1%
NYT Connections (extended)39%94.4%
ARC-AGI-193.5%75%
CritPt12%6.9%
Chess Puzzles30%31%
LMArena Hard Prompts15061480
Epoch Capabilities Index156.25152.92
ForecastBench60.361.2
EnigmaEval—18.2%
Thematic Generalization72.8%—
EBR-Bench19%—
Mystery Game Puzzles28%—
DTBench94.7%—
LMCA52.2%—

Math Claude Opus 4.7 leads

Claude Opus 4.7: 66.7 (#26), Gemini 3 Pro: 49.9 (#59)

Math benchmarks
BenchmarkClaude Opus 4.7Gemini 3 Pro
MathArena Final-Answer Competitions73.6%67%
OTIS Mock AIME 2024-202597.8%91.4%
ProofBench54%20%
LMArena Math14991476
FrontierMath (Feb 2025 set)43.8%37.6%
FrontierMath Tier 4 (v1)22.9%18.8%
FrontierMath (Tiers 1-3)70.2%—
FrontierMath Tier 431.7%—
Omni-MATH—55.5%

Knowledge Gemini 3 Pro leads

Claude Opus 4.7: 62.6 (#23), Gemini 3 Pro: 64.4 (#16)

Knowledge benchmarks
BenchmarkClaude Opus 4.7Gemini 3 Pro
GPQA Diamond90.2%92.6%
Humanity's Last Exam36.2%37.5%
Vectara Hallucination Rate12%13.6%
LMArena Expert15211475
SimpleQA Verified51.7%—
MMLU-Pro—90.3%
GPQA (HELM)—80.3%

Multimodal Gemini 3 Pro leads

Claude Opus 4.7: 41.2 (#38), Gemini 3 Pro: 57.6 (#2)

Multimodal benchmarks
BenchmarkClaude Opus 4.7Gemini 3 Pro
LMArena Vision13161305
LMArena Document14951434
GeoBench—84%
VPCT—91%
Blueprint-Bench 224.5%—
Furniture Assembly33.3%—

Multilingual Too close to call

Claude Opus 4.7: 57.3 (#10), Gemini 3 Pro: 56.9 (#16)

Multilingual benchmarks
BenchmarkClaude Opus 4.7Gemini 3 Pro
LMArena Non-English14801474
LMArena Chinese15311523
LMArena French15031492
LMArena German14951515
LMArena Japanese14721510
LMArena Korean14641448
LMArena Russian14941493
LMArena Spanish14951470

Instruction Following Claude Opus 4.7 leads

Claude Opus 4.7: 78.4 (#10), Gemini 3 Pro: 76.3 (#45)

Instruction Following benchmarks
BenchmarkClaude Opus 4.7Gemini 3 Pro
LMArena Instruction Following14981458
IFEval—87.7%

Long Context Claude Opus 4.7 leads

Claude Opus 4.7: 46.2 (#25), Gemini 3 Pro: 44.0 (#79)

Long Context benchmarks
BenchmarkClaude Opus 4.7Gemini 3 Pro
LMArena Longer Query15051471
CL-bench—15.8%

Writing & Preference Claude Opus 4.7 leads

Claude Opus 4.7: 75.1 (#8), Gemini 3 Pro: 66.4 (#35)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.7Gemini 3 Pro
LMArena Text14901479
LMArena Creative Writing14861482
EQ-Bench Creative Writing19141525
LMArena Multi-Turn15051484
WildBench—85.9%
EQ-Bench 41311—

Frequently asked questions

Is Claude Opus 4.7 better than Gemini 3 Pro?

Claude Opus 4.7 is the stronger model overall, scoring 58.3 to 54.8 on the Noometry Index.

Is Claude Opus 4.7 or Gemini 3 Pro better for coding?

Claude Opus 4.7 scores higher on coding benchmarks: 59.6 versus 51.6 in the Noometry coding category.

How many benchmarks do Claude Opus 4.7 and Gemini 3 Pro share?

46 benchmarks have published results for both models. Claude Opus 4.7 has 66 scored results on Noometry and Gemini 3 Pro has 67.

Related comparisons

Go deeper