Model comparison

GLM-5.3-Flash vs MiniMax-M3

GLM-5.3-Flash is the stronger model overall, scoring 51.8 to 43.8 on the Noometry Index.

Last verified . 31 shared benchmarks.

GLM-5.3-Flash Z.ai (Zhipu)

51.8

Rank #41 Confirmed

MiniMax-M3 MiniMax

43.8

Rank #85 Confirmed

Summary

  • They share 31 benchmarks with published results for both. GLM-5.3-Flash scores higher in 9 categories and MiniMax-M3 in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GLM-5.3-Flash leads 48.0 to 30.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 93.9% for GLM-5.3-Flash and 71.1% for MiniMax-M3.
  • GLM-5.3-Flash is cheaper at $0.15 / $0.50 per million input/output tokens, against $0.30 / $1.20 for MiniMax-M3.

Side by side

GLM-5.3-Flash and MiniMax-M3 specifications
GLM-5.3-FlashMiniMax-M3
ProviderZ.ai (Zhipu)MiniMax
Noometry Index51.843.8
Released2026-08-202026-06-01
WeightsOpenOpen
Context window1M1M
Max output131K512K
Input $ / M tokens$0.15$0.30
Output $ / M tokens$0.50$1.20
Results tracked4041

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.3-Flash leads

GLM-5.3-Flash: 53.1 (#31), MiniMax-M3: 41.8 (#118)

Coding benchmarks
BenchmarkGLM-5.3-FlashMiniMax-M3
FrontierCode31.8%14.7%
LMArena WebDev16091482
SciCode51.6%47.1%
LMArena Coding15081469
ALE-Bench303.55640.02
DeepSWE63.4%—
CursorBench36.8%—
FrontierSWE18.1%—

Agentic & Tool Use GLM-5.3-Flash leads

GLM-5.3-Flash: 34.2 (#47), MiniMax-M3: 22.6 (#130)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.3-FlashMiniMax-M3
APEX-Agents52.8%37.7%
OSWorld 2.0—4.6%
GBAEval—0.9%
GDP.pdf14%—
Vending-Bench 2—2,158

Reasoning GLM-5.3-Flash leads

GLM-5.3-Flash: 48.0 (#42), MiniMax-M3: 30.1 (#87)

Reasoning benchmarks
BenchmarkGLM-5.3-FlashMiniMax-M3
CritPt15.4%3.7%
Chess Puzzles14%14%
LMArena Hard Prompts14911447
Mystery Game Puzzles8%8%
Surface Evolver Bench52.5%55%
Epoch Capabilities Index151.88146.95
ARC-AGI-265.8%—
SimpleBench—45.8%
NYT Connections (extended)—65.1%
ARC-AGI-191%—
DTBench—78.9%
LMCA—33.7%
Bench to the Future 30.15—
ForecastBench—61.4

Math GLM-5.3-Flash leads

GLM-5.3-Flash: 53.3 (#47), MiniMax-M3: 40.0 (#95)

Math benchmarks
BenchmarkGLM-5.3-FlashMiniMax-M3
OTIS Mock AIME 2024-202593.9%71.1%
ProofBench21%18%
LMArena Math15001429
FrontierMath (Tiers 1-3)55.8%—
FrontierMath Tier 417.1%—

Knowledge Too close to call

GLM-5.3-Flash: 58.4 (#36), MiniMax-M3: 58.4 (#35)

Knowledge benchmarks
BenchmarkGLM-5.3-FlashMiniMax-M3
GPQA Diamond90.2%90.9%
LMArena Expert15131461

Multimodal GLM-5.3-Flash leads

GLM-5.3-Flash: 42.8 (#27), MiniMax-M3: 40.2 (#51)

Multimodal benchmarks
BenchmarkGLM-5.3-FlashMiniMax-M3
LMArena Vision12961253
LMArena Document—1435

Multilingual GLM-5.3-Flash leads

GLM-5.3-Flash: 56.0 (#25), MiniMax-M3: 53.0 (#75)

Multilingual benchmarks
BenchmarkGLM-5.3-FlashMiniMax-M3
LMArena Non-English14621420
LMArena Chinese15271463
LMArena French14961447
LMArena German14701426
LMArena Japanese14291381
LMArena Korean14461372
LMArena Russian14691428
LMArena Spanish14711432

Instruction Following GLM-5.3-Flash leads

GLM-5.3-Flash: 77.5 (#20), MiniMax-M3: 75.5 (#62)

Instruction Following benchmarks
BenchmarkGLM-5.3-FlashMiniMax-M3
LMArena Instruction Following14781433

Long Context GLM-5.3-Flash leads

GLM-5.3-Flash: 45.4 (#39), MiniMax-M3: 44.2 (#72)

Long Context benchmarks
BenchmarkGLM-5.3-FlashMiniMax-M3
LMArena Longer Query14821445

Writing & Preference GLM-5.3-Flash leads

GLM-5.3-Flash: 65.3 (#50), MiniMax-M3: 62.1 (#83)

Writing & Preference benchmarks
BenchmarkGLM-5.3-FlashMiniMax-M3
LMArena Text14711433
LMArena Creative Writing14421404
LMArena Multi-Turn14671442
EQ-Bench 4—1150

Frequently asked questions

Is GLM-5.3-Flash better than MiniMax-M3?

GLM-5.3-Flash is the stronger model overall, scoring 51.8 to 43.8 on the Noometry Index.

Which is cheaper, GLM-5.3-Flash or MiniMax-M3?

GLM-5.3-Flash is cheaper. It lists at $0.15 per million input tokens and $0.50 per million output tokens; MiniMax-M3 lists at $0.30 and $1.20.

Is GLM-5.3-Flash or MiniMax-M3 better for coding?

GLM-5.3-Flash scores higher on coding benchmarks: 53.1 versus 41.8 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do GLM-5.3-Flash and MiniMax-M3 share?

31 benchmarks have published results for both models. GLM-5.3-Flash has 40 scored results on Noometry and MiniMax-M3 has 41.

Related comparisons

Go deeper