Model comparison

Command A vs GLM-4.7-Flash

GLM-4.7-Flash is the stronger model overall, scoring 38.8 to 36.5 on the Noometry Index.

Last verified . 18 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Command A scores higher in 3 categories and GLM-4.7-Flash in 5 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GLM-4.7-Flash leads 40.6 to 27.2.
  • GLM-4.7-Flash is cheaper at $0.06 / $0.40 per million input/output tokens, against $2.50 / $10 for Command A.
  • Command A accepts more context: 256K tokens versus 200K.

Side by side

Command A and GLM-4.7-Flash specifications
Command AGLM-4.7-Flash
ProviderCohereZ.ai (Zhipu)
Noometry Index36.538.8
Released2025-03-132026-01-19
WeightsOpenOpen
Context window256K200K
Max output8K131K
Input $ / M tokens$2.50$0.06
Output $ / M tokens$10$0.40
Results tracked2421

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.7-Flash leads

Command A: 27.2 (#322), GLM-4.7-Flash: 40.6 (#135)

Coding benchmarks
BenchmarkCommand AGLM-4.7-Flash
LMArena Coding13301383
Aider Polyglot12%—

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), GLM-4.7-Flash: —

Agentic & Tool Use benchmarks
BenchmarkCommand AGLM-4.7-Flash
Berkeley Function Calling Leaderboard57.1%—

Reasoning GLM-4.7-Flash leads

Command A: 18.3 (#283), GLM-4.7-Flash: 20.9 (#229)

Reasoning benchmarks
BenchmarkCommand AGLM-4.7-Flash
LMArena Hard Prompts13261356
Kagi LLM Benchmark28.8%—
Chess Puzzles—0%
DTBench61.3%—
LMCA10.3%—

Math Too close to call

Command A: 36.2 (#171), GLM-4.7-Flash: 36.1 (#173)

Math benchmarks
BenchmarkCommand AGLM-4.7-Flash
LMArena Math13001355
OTIS Mock AIME 2024-2025—58.3%

Knowledge Command A leads

Command A: 37.1 (#159), GLM-4.7-Flash: 35.5 (#184)

Knowledge benchmarks
BenchmarkCommand AGLM-4.7-Flash
Vectara Hallucination Rate9.3%9.3%
LMArena Expert12951357
GPQA Diamond—60.5%

Multilingual GLM-4.7-Flash leads

Command A: 45.3 (#170), GLM-4.7-Flash: 46.5 (#158)

Multilingual benchmarks
BenchmarkCommand AGLM-4.7-Flash
LMArena Non-English13131330
LMArena Chinese13271403
LMArena French13511332
LMArena German13411337
LMArena Korean12851283
LMArena Russian13141332
LMArena Spanish13471350
LMArena Japanese1285—

Instruction Following Too close to call

Command A: 69.1 (#177), GLM-4.7-Flash: 70.1 (#167)

Instruction Following benchmarks
BenchmarkCommand AGLM-4.7-Flash
LMArena Instruction Following13091327

Long Context Too close to call

Command A: 40.6 (#151), GLM-4.7-Flash: 40.9 (#148)

Long Context benchmarks
BenchmarkCommand AGLM-4.7-Flash
LMArena Longer Query13341345

Writing & Preference Too close to call

Command A: 47.6 (#208), GLM-4.7-Flash: 47.4 (#210)

Writing & Preference benchmarks
BenchmarkCommand AGLM-4.7-Flash
LMArena Text13311351
LMArena Creative Writing13191297
EQ-Bench Creative Writing11451125
LMArena Multi-Turn13391342

Frequently asked questions

Is Command A better than GLM-4.7-Flash?

GLM-4.7-Flash is the stronger model overall, scoring 38.8 to 36.5 on the Noometry Index.

Which is cheaper, Command A or GLM-4.7-Flash?

GLM-4.7-Flash is cheaper. It lists at $0.06 per million input tokens and $0.40 per million output tokens; Command A lists at $2.50 and $10.

Is Command A or GLM-4.7-Flash better for coding?

GLM-4.7-Flash scores higher on coding benchmarks: 40.6 versus 27.2 in the Noometry coding category.

Which has the bigger context window?

Command A does, with 256K tokens against 200K.

How many benchmarks do Command A and GLM-4.7-Flash share?

18 benchmarks have published results for both models. Command A has 24 scored results on Noometry and GLM-4.7-Flash has 21.

Related comparisons

Go deeper