Model comparison

Claude Opus 4.1 vs Gemini 3.1 Pro Preview

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 41.0 on the Noometry Index.

Last verified . 43 shared benchmarks.

Claude Opus 4.1 Anthropic

41.0

Rank #142 Confirmed

Gemini 3.1 Pro Preview Google

56.7

Rank #23 Confirmed

Summary

  • They share 43 benchmarks with published results for both. Claude Opus 4.1 scores higher in 1 category and Gemini 3.1 Pro Preview in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemini 3.1 Pro Preview leads 62.1 to 22.3.
  • The biggest single-benchmark swing is Chess Puzzles: 7% for Claude Opus 4.1 and 55% for Gemini 3.1 Pro Preview.
  • Gemini 3.1 Pro Preview is cheaper at $2 / $12 per million input/output tokens, against $15 / $75 for Claude Opus 4.1.
  • Gemini 3.1 Pro Preview accepts more context: 1.05M tokens versus 200K.

Side by side

Claude Opus 4.1 and Gemini 3.1 Pro Preview specifications
Claude Opus 4.1Gemini 3.1 Pro Preview
ProviderAnthropicGoogle
Noometry Index41.056.7
Released2025-08-052026-02-19
WeightsProprietaryProprietary
Context window200K1.05M
Max output32K66K
Input $ / M tokens$15$2
Output $ / M tokens$75$12
Results tracked4871

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.1 leads

Claude Opus 4.1: 44.4 (#73), Gemini 3.1 Pro Preview: 42.5 (#99)

Coding benchmarks
BenchmarkClaude Opus 4.1Gemini 3.1 Pro Preview
SWE-bench Verified73.3%75.6%
LMArena WebDev13901447
WeirdML45.9%72.1%
LMArena Coding14791484
ALE-Bench674.771,161
AlgoTune1.342.02
DeepSWE—11.7%
SciCode—58.9%
GSO—22.6%
MirrorCode—8.9%

Agentic & Tool Use Gemini 3.1 Pro Preview leads

Claude Opus 4.1: 35.0 (#41), Gemini 3.1 Pro Preview: 37.7 (#34)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.1Gemini 3.1 Pro Preview
Terminal-Bench38%80.2%
DeepResearch Bench48.3%47.8%
LMArena Search11481211
METR Time Horizons66.8%77%
APEX-Agents—35.3%
GDPval43.6%—
τ²-bench Banking—26%
Cybench42%—
PostTrainBench—22%
BALROG—57%
ExploitBench—26.1%
GBAEval—0.8%
GDP.pdf—17%
Vending-Bench 2—3,774

Reasoning Gemini 3.1 Pro Preview leads

Claude Opus 4.1: 32.2 (#76), Gemini 3.1 Pro Preview: 71.7 (#12)

Reasoning benchmarks
BenchmarkClaude Opus 4.1Gemini 3.1 Pro Preview
SimpleBench60%79.6%
Chess Puzzles7%55%
EnigmaEval7.2%36.8%
EBR-Bench7.9%14.3%
LMArena Hard Prompts14431485
Mystery Game Puzzles21%34%
DTBench80%97.1%
LMCA37.1%53.8%
Epoch Capabilities Index144.12154.77
ForecastBench6259
ARC-AGI-2—77.1%
NYT Connections (extended)—97.4%
ARC-AGI-1—98%
CritPt—17.7%
Thematic Generalization—79.4%

Math Gemini 3.1 Pro Preview leads

Claude Opus 4.1: 22.3 (#277), Gemini 3.1 Pro Preview: 62.1 (#34)

Math benchmarks
BenchmarkClaude Opus 4.1Gemini 3.1 Pro Preview
FrontierMath (Tiers 1-3)12.6%59.6%
FrontierMath Tier 42.4%26.8%
OTIS Mock AIME 2024-202568.9%95.6%
LMArena Math14311485
FrontierMath (Feb 2025 set)7.2%36.9%
FrontierMath Tier 4 (v1)4.2%16.7%
MathArena Final-Answer Competitions—86.5%
ProofBench—26%

Knowledge Gemini 3.1 Pro Preview leads

Claude Opus 4.1: 42.0 (#101), Gemini 3.1 Pro Preview: 71.8 (#3)

Knowledge benchmarks
BenchmarkClaude Opus 4.1Gemini 3.1 Pro Preview
GPQA Diamond77.3%94.4%
Humanity's Last Exam11.5%46.4%
Vectara Hallucination Rate11.8%10.4%
LMArena Expert14391485
SimpleQA Verified—73.5%
Confabulations17.1%—

Multimodal Gemini 3.1 Pro Preview leads

Claude Opus 4.1: 26.8 (#119), Gemini 3.1 Pro Preview: 37.9 (#69)

Multimodal benchmarks
BenchmarkClaude Opus 4.1Gemini 3.1 Pro Preview
LMArena Vision—1296
VPCT35%—
Blueprint-Bench 2—26.5%
Furniture Assembly—26.7%
LMArena Document—1444

Multilingual Gemini 3.1 Pro Preview leads

Claude Opus 4.1: 52.0 (#95), Gemini 3.1 Pro Preview: 57.0 (#12)

Multilingual benchmarks
BenchmarkClaude Opus 4.1Gemini 3.1 Pro Preview
LMArena Non-English14051477
LMArena Chinese14271529
LMArena French14311487
LMArena German14131491
LMArena Japanese13781493
LMArena Korean13801455
LMArena Russian14221498
LMArena Spanish14481479

Instruction Following Gemini 3.1 Pro Preview leads

Claude Opus 4.1: 75.6 (#58), Gemini 3.1 Pro Preview: 77.0 (#32)

Instruction Following benchmarks
BenchmarkClaude Opus 4.1Gemini 3.1 Pro Preview
LMArena Instruction Following14351466

Long Context Gemini 3.1 Pro Preview leads

Claude Opus 4.1: 44.5 (#63), Gemini 3.1 Pro Preview: 47.4 (#18)

Long Context benchmarks
BenchmarkClaude Opus 4.1Gemini 3.1 Pro Preview
LMArena Longer Query14551483
CL-bench—20.8%
CL-bench Life—16.9%

Writing & Preference Gemini 3.1 Pro Preview leads

Claude Opus 4.1: 62.4 (#74), Gemini 3.1 Pro Preview: 66.1 (#37)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.1Gemini 3.1 Pro Preview
LMArena Text14191481
LMArena Creative Writing14121482
LMArena Multi-Turn14441488
Short-Story Creative Writing84.7%—
EQ-Bench Creative Writing—1491
EQ-Bench 4—1142

Frequently asked questions

Is Claude Opus 4.1 better than Gemini 3.1 Pro Preview?

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 41.0 on the Noometry Index.

Which is cheaper, Claude Opus 4.1 or Gemini 3.1 Pro Preview?

Gemini 3.1 Pro Preview is cheaper. It lists at $2 per million input tokens and $12 per million output tokens; Claude Opus 4.1 lists at $15 and $75.

Is Claude Opus 4.1 or Gemini 3.1 Pro Preview better for coding?

Claude Opus 4.1 scores higher on coding benchmarks: 44.4 versus 42.5 in the Noometry coding category.

Which has the bigger context window?

Gemini 3.1 Pro Preview does, with 1.05M tokens against 200K.

How many benchmarks do Claude Opus 4.1 and Gemini 3.1 Pro Preview share?

43 benchmarks have published results for both models. Claude Opus 4.1 has 48 scored results on Noometry and Gemini 3.1 Pro Preview has 71.

Related comparisons

Go deeper