Model comparison

Claude Opus 4 vs Gemma 4 31B IT

Claude Opus 4 and Gemma 4 31B IT score almost the same on the Noometry Index (43.1 vs 43.5), so choose on price, context window or the category you care about most.

Last verified . 25 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

Gemma 4 31B IT Google

43.5

Rank #90 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude Opus 4 scores higher in 5 categories and Gemma 4 31B IT in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in multimodal, where Gemma 4 31B IT leads 41.6 to 31.5.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 74.3% for Claude Opus 4 and 63.5% for Gemma 4 31B IT.
  • Gemma 4 31B IT is cheaper at $0.09 / $0.34 per million input/output tokens, against $15 / $75 for Claude Opus 4.
  • Gemma 4 31B IT accepts more context: 262K tokens versus 200K.
  • Gemma 4 31B IT has downloadable open weights; the other is API-only.

Side by side

Claude Opus 4 and Gemma 4 31B IT specifications
Claude Opus 4Gemma 4 31B IT
ProviderAnthropicGoogle
Noometry Index43.143.5
Released2025-05-222026-04-02
WeightsProprietaryOpen
Context window200K262K
Max output32K33K
Input $ / M tokens$15$0.09
Output $ / M tokens$75$0.34
Results tracked5635

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4 leads

Claude Opus 4: 47.2 (#62), Gemma 4 31B IT: 42.3 (#108)

Coding benchmarks
BenchmarkClaude Opus 4Gemma 4 31B IT
WeirdML43.7%52.3%
LMArena Coding14421459
SWE-bench Verified70.7%—
SWE-bench Verified (bash only)67.6%—
Aider Polyglot72%—
LMArena WebDev—1366
SciCode—43.4%
GSO6.9%—
ALE-Bench—925.5
AlgoTune1.33—

Agentic & Tool Use Not comparable

Claude Opus 4: 34.8 (#42), Gemma 4 31B IT: —

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4Gemma 4 31B IT
Cybench38%—
DeepResearch Bench46.8%—
LMArena Search1127—
METR Time Horizons63.9%—

Reasoning Too close to call

Claude Opus 4: 27.3 (#121), Gemma 4 31B IT: 27.2 (#122)

Reasoning benchmarks
BenchmarkClaude Opus 4Gemma 4 31B IT
Kagi LLM Benchmark74.3%63.5%
CritPt0.3%1.4%
LMArena Hard Prompts13991448
DTBench81.6%82.7%
LMCA37.4%39.3%
Epoch Capabilities Index142.67142.74
ARC-AGI-28.6%—
SimpleBench58.8%—
NYT Connections (extended)—70.6%
ARC-AGI-135.7%—
Chess Puzzles—5%
EnigmaEval5.6%—
Thematic Generalization—53%
Surface Evolver Bench—30.6%
ForecastBench61.1—

Math Gemma 4 31B IT leads

Claude Opus 4: 42.0 (#86), Gemma 4 31B IT: 43.2 (#81)

Math benchmarks
BenchmarkClaude Opus 4Gemma 4 31B IT
OTIS Mock AIME 2024-202564.4%73.3%
LMArena Math13901465
Omni-MATH61.6%—
MATH Level 585%—
FrontierMath (Feb 2025 set)4.5%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Claude Opus 4 leads

Claude Opus 4: 44.0 (#88), Gemma 4 31B IT: 37.9 (#151)

Knowledge benchmarks
BenchmarkClaude Opus 4Gemma 4 31B IT
GPQA Diamond76.3%75.8%
Vectara Hallucination Rate12%7.4%
LMArena Expert13861465
Humanity's Last Exam10.7%—
SimpleQA Verified—10.4%
MMLU-Pro87.5%—
Confabulations15.9%—
GPQA (HELM)70.8%—

Multimodal Gemma 4 31B IT leads

Claude Opus 4: 31.5 (#106), Gemma 4 31B IT: 41.6 (#34)

Multimodal benchmarks
BenchmarkClaude Opus 4Gemma 4 31B IT
LMArena Vision11921277
GeoBench49%—
VPCT38%—
LMArena Document—1425

Multilingual Gemma 4 31B IT leads

Claude Opus 4: 48.8 (#138), Gemma 4 31B IT: 53.8 (#57)

Multilingual benchmarks
BenchmarkClaude Opus 4Gemma 4 31B IT
LMArena Non-English13621431
LMArena Chinese13861476
LMArena French13721435
LMArena Russian13921460
LMArena Spanish13891444
LMArena German1391—
LMArena Japanese1331—
LMArena Korean1321—

Instruction Following Claude Opus 4 leads

Claude Opus 4: 77.1 (#28), Gemma 4 31B IT: 75.5 (#61)

Instruction Following benchmarks
BenchmarkClaude Opus 4Gemma 4 31B IT
LMArena Instruction Following14061433
IFEval91.8%—

Long Context Gemma 4 31B IT leads

Claude Opus 4: 39.6 (#172), Gemma 4 31B IT: 44.2 (#71)

Long Context benchmarks
BenchmarkClaude Opus 4Gemma 4 31B IT
LMArena Longer Query14221446
Fiction.LiveBench61.1%—

Writing & Preference Too close to call

Claude Opus 4: 61.2 (#89), Gemma 4 31B IT: 60.5 (#96)

Writing & Preference benchmarks
BenchmarkClaude Opus 4Gemma 4 31B IT
LMArena Text13771443
LMArena Creative Writing13871415
EQ-Bench Creative Writing15801368
LMArena Multi-Turn13961452
Short-Story Creative Writing83.6%—
WildBench85.2%—
EQ-Bench 4—1120

Frequently asked questions

Is Claude Opus 4 better than Gemma 4 31B IT?

Claude Opus 4 and Gemma 4 31B IT score almost the same on the Noometry Index (43.1 vs 43.5), so choose on price, context window or the category you care about most.

Which is cheaper, Claude Opus 4 or Gemma 4 31B IT?

Gemma 4 31B IT is cheaper. It lists at $0.09 per million input tokens and $0.34 per million output tokens; Claude Opus 4 lists at $15 and $75.

Is Claude Opus 4 or Gemma 4 31B IT better for coding?

Claude Opus 4 scores higher on coding benchmarks: 47.2 versus 42.3 in the Noometry coding category.

Which has the bigger context window?

Gemma 4 31B IT does, with 262K tokens against 200K.

How many benchmarks do Claude Opus 4 and Gemma 4 31B IT share?

25 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and Gemma 4 31B IT has 35.

Related comparisons

Go deeper