Model comparison

GLM-4.5V vs Olmo 3.1 32b Instruct

GLM-4.5V and Olmo 3.1 32b Instruct score almost the same on the Noometry Index (39.8 vs 39.4), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

GLM-4.5V Z.ai (Zhipu)

39.8

Rank #158 Confirmed

Summary

  • They share 13 benchmarks with published results for both. GLM-4.5V scores higher in 7 categories and Olmo 3.1 32b Instruct in 1 category; 4 gaps are clear of the uncertainty.

Side by side

GLM-4.5V and Olmo 3.1 32b Instruct specifications
GLM-4.5VOlmo 3.1 32b Instruct
ProviderZ.ai (Zhipu)Allen Institute for AI (Ai2)
Noometry Index39.839.4
Released2025-08-11—
WeightsOpenOpen
Context window64K—
Max output16K—
Input $ / M tokens$0.60—
Output $ / M tokens$1.80—
Results tracked1516

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GLM-4.5V: 39.5 (#155), Olmo 3.1 32b Instruct: 39.5 (#157)

Coding benchmarks
BenchmarkGLM-4.5VOlmo 3.1 32b Instruct
LMArena Coding13471347

Reasoning Too close to call

GLM-4.5V: 27.4 (#119), Olmo 3.1 32b Instruct: 26.4 (#132)

Reasoning benchmarks
BenchmarkGLM-4.5VOlmo 3.1 32b Instruct
LMArena Hard Prompts13341322
Kagi LLM Benchmark59.8%—

Math GLM-4.5V leads

GLM-4.5V: 37.4 (#159), Olmo 3.1 32b Instruct: 36.3 (#167)

Math benchmarks
BenchmarkGLM-4.5VOlmo 3.1 32b Instruct
LMArena Math13541305

Knowledge GLM-4.5V leads

GLM-4.5V: 37.5 (#156), Olmo 3.1 32b Instruct: 36.1 (#175)

Knowledge benchmarks
BenchmarkGLM-4.5VOlmo 3.1 32b Instruct
LMArena Expert13531308

Multimodal Not comparable

GLM-4.5V: 34.3 (#92), Olmo 3.1 32b Instruct: —

Multimodal benchmarks
BenchmarkGLM-4.5VOlmo 3.1 32b Instruct
LMArena Vision1154—

Multilingual GLM-4.5V leads

GLM-4.5V: 44.6 (#177), Olmo 3.1 32b Instruct: 42.6 (#191)

Multilingual benchmarks
BenchmarkGLM-4.5VOlmo 3.1 32b Instruct
LMArena Non-English13031275
LMArena Chinese13371304
LMArena Russian12981268
LMArena Spanish13361336
LMArena French—1328
LMArena German—1282
LMArena Korean—1206

Instruction Following Too close to call

GLM-4.5V: 69.2 (#175), Olmo 3.1 32b Instruct: 68.6 (#187)

Instruction Following benchmarks
BenchmarkGLM-4.5VOlmo 3.1 32b Instruct
LMArena Instruction Following13111299

Long Context Too close to call

GLM-4.5V: 39.6 (#171), Olmo 3.1 32b Instruct: 39.9 (#166)

Long Context benchmarks
BenchmarkGLM-4.5VOlmo 3.1 32b Instruct
LMArena Longer Query13041312

Writing & Preference GLM-4.5V leads

GLM-4.5V: 52.5 (#170), Olmo 3.1 32b Instruct: 50.2 (#185)

Writing & Preference benchmarks
BenchmarkGLM-4.5VOlmo 3.1 32b Instruct
LMArena Text13331311
LMArena Creative Writing12951264
LMArena Multi-Turn13321309

Frequently asked questions

Is GLM-4.5V better than Olmo 3.1 32b Instruct?

GLM-4.5V and Olmo 3.1 32b Instruct score almost the same on the Noometry Index (39.8 vs 39.4), so choose on price, context window or the category you care about most.

Is GLM-4.5V or Olmo 3.1 32b Instruct better for coding?

They score almost the same on coding (39.5 vs 39.5); test both on your own repository before choosing.

How many benchmarks do GLM-4.5V and Olmo 3.1 32b Instruct share?

13 benchmarks have published results for both models. GLM-4.5V has 15 scored results on Noometry and Olmo 3.1 32b Instruct has 16.

Related comparisons

Go deeper