Model comparison

GLM-4.5-Air vs Step 3.5 Flash

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 38.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • They share 17 benchmarks with published results for both. GLM-4.5-Air scores higher in 1 category and Step 3.5 Flash in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Step 3.5 Flash leads 42.4 to 33.3.
  • Step 3.5 Flash is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.20 / $1.10 for GLM-4.5-Air.
  • Step 3.5 Flash accepts more context: 256K tokens versus 131K.

Side by side

GLM-4.5-Air and Step 3.5 Flash specifications
GLM-4.5-AirStep 3.5 Flash
ProviderZ.ai (Zhipu)StepFun
Noometry Index38.942.3
Released2025-07-202026-01-29
WeightsOpenOpen
Context window131K256K
Max output98K256K
Input $ / M tokens$0.20$0.10
Output $ / M tokens$1.10$0.30
Results tracked2719

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3.5 Flash leads

GLM-4.5-Air: 33.3 (#259), Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkGLM-4.5-AirStep 3.5 Flash
LMArena Coding13971436
GSO2.9%—

Reasoning GLM-4.5-Air leads

GLM-4.5-Air: 24.1 (#166), Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkGLM-4.5-AirStep 3.5 Flash
LMArena Hard Prompts13791411
Kagi LLM Benchmark43%—
NYT Connections (extended)—28.4%
ForecastBench59.2—

Math Step 3.5 Flash leads

GLM-4.5-Air: 36.2 (#170), Step 3.5 Flash: 42.6 (#84)

Math benchmarks
BenchmarkGLM-4.5-AirStep 3.5 Flash
LMArena Math13961408
MathArena Final-Answer Competitions—66.8%
Omni-MATH39.1%—

Knowledge Step 3.5 Flash leads

GLM-4.5-Air: 35.0 (#191), Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkGLM-4.5-AirStep 3.5 Flash
LMArena Expert13701421
Humanity's Last Exam8.1%—
MMLU-Pro76.2%—
Vectara Hallucination Rate9.3%—
GPQA (HELM)59.4%—

Multilingual Step 3.5 Flash leads

GLM-4.5-Air: 49.1 (#135), Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkGLM-4.5-AirStep 3.5 Flash
LMArena Non-English13661385
LMArena Chinese14261447
LMArena French13991421
LMArena German13771405
LMArena Japanese13481354
LMArena Korean13081352
LMArena Russian13731385
LMArena Spanish13861419

Instruction Following Step 3.5 Flash leads

GLM-4.5-Air: 69.6 (#171), Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkGLM-4.5-AirStep 3.5 Flash
LMArena Instruction Following13541385
IFEval81.2%—

Long Context Step 3.5 Flash leads

GLM-4.5-Air: 41.6 (#135), Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkGLM-4.5-AirStep 3.5 Flash
LMArena Longer Query13661402

Writing & Preference Step 3.5 Flash leads

GLM-4.5-Air: 55.9 (#139), Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkGLM-4.5-AirStep 3.5 Flash
LMArena Text13841403
LMArena Creative Writing13431357
LMArena Multi-Turn13711405
WildBench78.9%—

Frequently asked questions

Is GLM-4.5-Air better than Step 3.5 Flash?

Step 3.5 Flash is the stronger model overall, scoring 42.3 to 38.9 on the Noometry Index.

Which is cheaper, GLM-4.5-Air or Step 3.5 Flash?

Step 3.5 Flash is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; GLM-4.5-Air lists at $0.20 and $1.10.

Is GLM-4.5-Air or Step 3.5 Flash better for coding?

Step 3.5 Flash scores higher on coding benchmarks: 42.4 versus 33.3 in the Noometry coding category.

Which has the bigger context window?

Step 3.5 Flash does, with 256K tokens against 131K.

How many benchmarks do GLM-4.5-Air and Step 3.5 Flash share?

17 benchmarks have published results for both models. GLM-4.5-Air has 27 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper