Model comparison

Chatgpt 4o Latest 20250326 vs Gemma 4 31B IT

Chatgpt 4o Latest 20250326 and Gemma 4 31B IT score almost the same on the Noometry Index (43.8 vs 43.5), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Gemma 4 31B IT Google

43.5

Rank #90 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 3 categories and Gemma 4 31B IT in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Chatgpt 4o Latest 20250326 leads 33.8 to 27.2.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 75% for Chatgpt 4o Latest 20250326 and 63.5% for Gemma 4 31B IT.
  • Gemma 4 31B IT has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and Gemma 4 31B IT specifications
Chatgpt 4o Latest 20250326Gemma 4 31B IT
ProviderOpenAIGoogle
Noometry Index43.843.5
Released—2026-04-02
WeightsProprietaryOpen
Context window—262K
Max output—33K
Input $ / M tokens—$0.09
Output $ / M tokens—$0.34
Results tracked2135

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Chatgpt 4o Latest 20250326: 41.6 (#122), Gemma 4 31B IT: 42.3 (#108)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemma 4 31B IT
LMArena Coding14131459
LMArena WebDev—1366
SciCode—43.4%
WeirdML—52.3%
ALE-Bench—925.5

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Gemma 4 31B IT: 27.2 (#122)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemma 4 31B IT
Kagi LLM Benchmark75%63.5%
LMArena Hard Prompts14241448
NYT Connections (extended)—70.6%
CritPt—1.4%
Chess Puzzles—5%
Thematic Generalization—53%
DTBench—82.7%
LMCA—39.3%
Surface Evolver Bench—30.6%
Epoch Capabilities Index—142.74

Math Gemma 4 31B IT leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Gemma 4 31B IT: 43.2 (#81)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemma 4 31B IT
LMArena Math14071465
OTIS Mock AIME 2024-2025—73.3%

Knowledge Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Gemma 4 31B IT: 37.9 (#151)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemma 4 31B IT
LMArena Expert14011465
GPQA Diamond—75.8%
SimpleQA Verified—10.4%
Confabulations16.6%—
Vectara Hallucination Rate—7.4%

Multimodal Gemma 4 31B IT leads

Chatgpt 4o Latest 20250326: 39.6 (#58), Gemma 4 31B IT: 41.6 (#34)

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemma 4 31B IT
LMArena Vision12431277
LMArena Document—1425

Multilingual Too close to call

Chatgpt 4o Latest 20250326: 52.9 (#76), Gemma 4 31B IT: 53.8 (#57)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemma 4 31B IT
LMArena Non-English14191431
LMArena Chinese14571476
LMArena French14461435
LMArena Russian14291460
LMArena Spanish14341444
LMArena German1423—
LMArena Japanese1405—
LMArena Korean1396—

Instruction Following Gemma 4 31B IT leads

Chatgpt 4o Latest 20250326: 74.0 (#107), Gemma 4 31B IT: 75.5 (#61)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemma 4 31B IT
LMArena Instruction Following14031433

Long Context Gemma 4 31B IT leads

Chatgpt 4o Latest 20250326: 43.1 (#107), Gemma 4 31B IT: 44.2 (#71)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemma 4 31B IT
LMArena Longer Query14131446

Writing & Preference Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Gemma 4 31B IT: 60.5 (#96)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemma 4 31B IT
LMArena Text14291443
LMArena Creative Writing14051415
EQ-Bench Creative Writing15011368
LMArena Multi-Turn14541452
EQ-Bench 4—1120

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Gemma 4 31B IT?

Chatgpt 4o Latest 20250326 and Gemma 4 31B IT score almost the same on the Noometry Index (43.8 vs 43.5), so choose on price, context window or the category you care about most.

Is Chatgpt 4o Latest 20250326 or Gemma 4 31B IT better for coding?

They score almost the same on coding (41.6 vs 42.3); test both on your own repository before choosing.

How many benchmarks do Chatgpt 4o Latest 20250326 and Gemma 4 31B IT share?

17 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Gemma 4 31B IT has 35.

Related comparisons

Go deeper