Model comparison

GLM-4.7-Flash vs Olmo 3.1 32b Instruct

GLM-4.7-Flash and Olmo 3.1 32b Instruct score almost the same on the Noometry Index (38.8 vs 39.4), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

Summary

  • They share 16 benchmarks with published results for both. GLM-4.7-Flash scores higher in 4 categories and Olmo 3.1 32b Instruct in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Olmo 3.1 32b Instruct leads 26.4 to 20.9.

Side by side

GLM-4.7-Flash and Olmo 3.1 32b Instruct specifications
GLM-4.7-FlashOlmo 3.1 32b Instruct
ProviderZ.ai (Zhipu)Allen Institute for AI (Ai2)
Noometry Index38.839.4
Released2026-01-19—
WeightsOpenOpen
Context window200K—
Max output131K—
Input $ / M tokens$0.06—
Output $ / M tokens$0.40—
Results tracked2116

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.7-Flash leads

GLM-4.7-Flash: 40.6 (#135), Olmo 3.1 32b Instruct: 39.5 (#157)

Coding benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Instruct
LMArena Coding13831347

Reasoning Olmo 3.1 32b Instruct leads

GLM-4.7-Flash: 20.9 (#229), Olmo 3.1 32b Instruct: 26.4 (#132)

Reasoning benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Instruct
LMArena Hard Prompts13561322
Chess Puzzles0%—

Math Too close to call

GLM-4.7-Flash: 36.1 (#173), Olmo 3.1 32b Instruct: 36.3 (#167)

Math benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Instruct
LMArena Math13551305
OTIS Mock AIME 2024-202558.3%—

Knowledge Too close to call

GLM-4.7-Flash: 35.5 (#184), Olmo 3.1 32b Instruct: 36.1 (#175)

Knowledge benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Instruct
LMArena Expert13571308
GPQA Diamond60.5%—
Vectara Hallucination Rate9.3%—

Multilingual GLM-4.7-Flash leads

GLM-4.7-Flash: 46.5 (#158), Olmo 3.1 32b Instruct: 42.6 (#191)

Multilingual benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Instruct
LMArena Non-English13301275
LMArena Chinese14031304
LMArena French13321328
LMArena German13371282
LMArena Korean12831206
LMArena Russian13321268
LMArena Spanish13501336

Instruction Following GLM-4.7-Flash leads

GLM-4.7-Flash: 70.1 (#167), Olmo 3.1 32b Instruct: 68.6 (#187)

Instruction Following benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Instruct
LMArena Instruction Following13271299

Long Context GLM-4.7-Flash leads

GLM-4.7-Flash: 40.9 (#148), Olmo 3.1 32b Instruct: 39.9 (#166)

Long Context benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Instruct
LMArena Longer Query13451312

Writing & Preference Olmo 3.1 32b Instruct leads

GLM-4.7-Flash: 47.4 (#210), Olmo 3.1 32b Instruct: 50.2 (#185)

Writing & Preference benchmarks
BenchmarkGLM-4.7-FlashOlmo 3.1 32b Instruct
LMArena Text13511311
LMArena Creative Writing12971264
LMArena Multi-Turn13421309
EQ-Bench Creative Writing1125—

Frequently asked questions

Is GLM-4.7-Flash better than Olmo 3.1 32b Instruct?

GLM-4.7-Flash and Olmo 3.1 32b Instruct score almost the same on the Noometry Index (38.8 vs 39.4), so choose on price, context window or the category you care about most.

Is GLM-4.7-Flash or Olmo 3.1 32b Instruct better for coding?

GLM-4.7-Flash scores higher on coding benchmarks: 40.6 versus 39.5 in the Noometry coding category.

How many benchmarks do GLM-4.7-Flash and Olmo 3.1 32b Instruct share?

16 benchmarks have published results for both models. GLM-4.7-Flash has 21 scored results on Noometry and Olmo 3.1 32b Instruct has 16.

Related comparisons

Go deeper