Model comparison

Gemma 4 31B IT vs GPT-5.1

GPT-5.1 is the stronger model overall, scoring 49.0 to 43.5 on the Noometry Index. Gemma 4 31B IT costs 23× less per token, which makes it the better buy when GPT-5.1's lead doesn't matter for your workload.

Last verified . 29 shared benchmarks.

Gemma 4 31B IT Google

43.5

Rank #90 Confirmed

GPT-5.1 OpenAI

49.0

Rank #53 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Gemma 4 31B IT scores higher in 0 categories and GPT-5.1 in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-5.1 leads 50.6 to 37.9.
  • The biggest single-benchmark swing is SimpleQA Verified: 10.4% for Gemma 4 31B IT and 48% for GPT-5.1.
  • Gemma 4 31B IT is cheaper at $0.09 / $0.34 per million input/output tokens, against $1.25 / $10 for GPT-5.1.
  • GPT-5.1 accepts more context: 400K tokens versus 262K.
  • Gemma 4 31B IT has downloadable open weights; the other is API-only.

Side by side

Gemma 4 31B IT and GPT-5.1 specifications
Gemma 4 31B ITGPT-5.1
ProviderGoogleOpenAI
Noometry Index43.549.0
Released2026-04-022025-11-13
WeightsOpenProprietary
Context window262K400K
Max output33K128K
Input $ / M tokens$0.09$1.25
Output $ / M tokens$0.34$10
Results tracked3563

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.1 leads

Gemma 4 31B IT: 42.3 (#108), GPT-5.1: 46.4 (#66)

Coding benchmarks
BenchmarkGemma 4 31B ITGPT-5.1
LMArena WebDev13661395
SciCode43.4%43.3%
WeirdML52.3%60.8%
LMArena Coding14591454
ALE-Bench925.51,192
SWE-bench Verified—68%
SWE-bench Verified (bash only)—66%
GSO—13.7%
LiveBench Coding—72.5%

Agentic & Tool Use Not comparable

Gemma 4 31B IT: —, GPT-5.1: 32.7 (#60)

Agentic & Tool Use benchmarks
BenchmarkGemma 4 31B ITGPT-5.1
Terminal-Bench—47.6%
DeepResearch Bench—42.8%
LMArena Search—1199
Vending-Bench 2—1,473

Reasoning GPT-5.1 leads

Gemma 4 31B IT: 27.2 (#122), GPT-5.1: 39.8 (#58)

Reasoning benchmarks
BenchmarkGemma 4 31B ITGPT-5.1
CritPt1.4%4.9%
Chess Puzzles5%32%
LMArena Hard Prompts14481457
DTBench82.7%90.1%
LMCA39.3%43.9%
Epoch Capabilities Index142.74149.64
ARC-AGI-2—17.6%
SimpleBench—53.2%
Kagi LLM Benchmark63.5%—
NYT Connections (extended)70.6%—
ARC-AGI-1—72.8%
EnigmaEval—11.2%
Thematic Generalization53%—
LiveBench Reasoning—95.8%
Mystery Game Puzzles—19%
LiveBench Data Analysis—72.1%
Surface Evolver Bench30.6%—
ForecastBench—58.1
LiveBench—78.8%

Math GPT-5.1 leads

Gemma 4 31B IT: 43.2 (#81), GPT-5.1: 52.2 (#51)

Math benchmarks
BenchmarkGemma 4 31B ITGPT-5.1
OTIS Mock AIME 2024-202573.3%88.6%
LMArena Math14651447
Omni-MATH—46.4%
LiveBench Math—94.5%
FrontierMath (Feb 2025 set)—31%
FrontierMath Tier 4 (v1)—12.5%

Knowledge GPT-5.1 leads

Gemma 4 31B IT: 37.9 (#151), GPT-5.1: 50.6 (#71)

Knowledge benchmarks
BenchmarkGemma 4 31B ITGPT-5.1
GPQA Diamond75.8%87.6%
SimpleQA Verified10.4%48%
Vectara Hallucination Rate7.4%10.9%
LMArena Expert14651470
Humanity's Last Exam—23.7%
MMLU-Pro—57.9%
GPQA (HELM)—44.2%

Multimodal GPT-5.1 leads

Gemma 4 31B IT: 41.6 (#34), GPT-5.1: 44.8 (#19)

Multimodal benchmarks
BenchmarkGemma 4 31B ITGPT-5.1
LMArena Vision12771250
LMArena Document14251403
VPCT—58.7%

Multilingual Too close to call

Gemma 4 31B IT: 53.8 (#57), GPT-5.1: 53.8 (#56)

Multilingual benchmarks
BenchmarkGemma 4 31B ITGPT-5.1
LMArena Non-English14311431
LMArena Chinese14761495
LMArena French14351450
LMArena Russian14601435
LMArena Spanish14441433
LMArena German—1438
LMArena Japanese—1453
LMArena Korean—1401

Instruction Following GPT-5.1 leads

Gemma 4 31B IT: 75.5 (#61), GPT-5.1: 83.9 (#1)

Instruction Following benchmarks
BenchmarkGemma 4 31B ITGPT-5.1
LMArena Instruction Following14331443
LiveBench Instruction Following—93.3%
IFEval—93.5%

Long Context GPT-5.1 leads

Gemma 4 31B IT: 44.2 (#71), GPT-5.1: 47.6 (#14)

Long Context benchmarks
BenchmarkGemma 4 31B ITGPT-5.1
LMArena Longer Query14461447
CL-bench—23.7%
CL-bench Life—17.3%

Writing & Preference GPT-5.1 leads

Gemma 4 31B IT: 60.5 (#96), GPT-5.1: 64.5 (#55)

Writing & Preference benchmarks
BenchmarkGemma 4 31B ITGPT-5.1
LMArena Text14431443
LMArena Creative Writing14151427
LMArena Multi-Turn14521450
EQ-Bench Creative Writing1368—
WildBench—86.3%
EQ-Bench 41120—
LiveBench Language—80.2%

Frequently asked questions

Is Gemma 4 31B IT better than GPT-5.1?

GPT-5.1 is the stronger model overall, scoring 49.0 to 43.5 on the Noometry Index. Gemma 4 31B IT costs 23× less per token, which makes it the better buy when GPT-5.1's lead doesn't matter for your workload.

Which is cheaper, Gemma 4 31B IT or GPT-5.1?

Gemma 4 31B IT is cheaper. It lists at $0.09 per million input tokens and $0.34 per million output tokens; GPT-5.1 lists at $1.25 and $10.

Is Gemma 4 31B IT or GPT-5.1 better for coding?

GPT-5.1 scores higher on coding benchmarks: 46.4 versus 42.3 in the Noometry coding category.

Which has the bigger context window?

GPT-5.1 does, with 400K tokens against 262K.

How many benchmarks do Gemma 4 31B IT and GPT-5.1 share?

29 benchmarks have published results for both models. Gemma 4 31B IT has 35 scored results on Noometry and GPT-5.1 has 63.

Related comparisons

Go deeper