Model comparison

GLM-4.7-Flash vs Olmo 3.1 32b Think

GLM-4.7-Flash and Olmo 3.1 32b Think score almost the same on the Noometry Index (38.8 vs 37.9), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Summary

  • They share 15 benchmarks with published results for both. GLM-4.7-Flash scores higher in 5 categories and Olmo 3.1 32b Think in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where GLM-4.7-Flash leads 46.5 to 38.1.

Side by side

GLM-4.7-Flash and Olmo 3.1 32b Think specifications
GLM-4.7-FlashOlmo 3.1 32b Think
ProviderZ.ai (Zhipu)Allen Institute for AI (Ai2)
Noometry Index38.837.9
Released2026-01-19—
WeightsOpenOpen
Context window200K—
Max output131K—
Input $ / M tokens$0.06—
Output $ / M tokens$0.40—
Results tracked2115

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.7-Flash leads

GLM-4.7-Flash: 40.6 (#135), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Think
LMArena Coding13831291

Reasoning Olmo 3.1 32b Think leads

GLM-4.7-Flash: 20.9 (#229), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Think
LMArena Hard Prompts13561272
Chess Puzzles0%—

Math Too close to call

GLM-4.7-Flash: 36.1 (#173), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Think
LMArena Math13551305
OTIS Mock AIME 2024-202558.3%—

Knowledge Too close to call

GLM-4.7-Flash: 35.5 (#184), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Think
LMArena Expert13571295
GPQA Diamond60.5%—
Vectara Hallucination Rate9.3%—

Multilingual GLM-4.7-Flash leads

GLM-4.7-Flash: 46.5 (#158), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Think
LMArena Non-English13301209
LMArena Chinese14031242
LMArena French13321260
LMArena German13371262
LMArena Russian13321193
LMArena Spanish13501289
LMArena Korean1283—

Instruction Following GLM-4.7-Flash leads

GLM-4.7-Flash: 70.1 (#167), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Think
LMArena Instruction Following13271247

Long Context GLM-4.7-Flash leads

GLM-4.7-Flash: 40.9 (#148), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Think
LMArena Longer Query13451272

Writing & Preference GLM-4.7-Flash leads

GLM-4.7-Flash: 47.4 (#210), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Think
LMArena Text13511272
LMArena Creative Writing12971226
LMArena Multi-Turn13421252
EQ-Bench Creative Writing1125—

Frequently asked questions

Is GLM-4.7-Flash better than Olmo 3.1 32b Think?

GLM-4.7-Flash and Olmo 3.1 32b Think score almost the same on the Noometry Index (38.8 vs 37.9), so choose on price, context window or the category you care about most.

Is GLM-4.7-Flash or Olmo 3.1 32b Think better for coding?

GLM-4.7-Flash scores higher on coding benchmarks: 40.6 versus 37.7 in the Noometry coding category.

How many benchmarks do GLM-4.7-Flash and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. GLM-4.7-Flash has 21 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper