Model comparison

GLM-4.5-Air vs Mistral Large

GLM-4.5-Air is the stronger model overall, scoring 38.9 to 31.9 on the Noometry Index.

Last verified . 24 shared benchmarks.

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 24 benchmarks with published results for both. GLM-4.5-Air scores higher in 7 categories and Mistral Large in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-4.5-Air leads 36.2 to 18.2.
  • The biggest single-benchmark swing is MMLU-Pro: 76.2% for GLM-4.5-Air and 59.9% for Mistral Large.
  • GLM-4.5-Air is cheaper at $0.20 / $1.10 per million input/output tokens, against $2 / $6 for Mistral Large.

Side by side

GLM-4.5-Air and Mistral Large specifications
GLM-4.5-AirMistral Large
ProviderZ.ai (Zhipu)Mistral AI
Noometry Index38.931.9
Released2025-07-202024-02-26
WeightsOpenOpen
Context window131K131K
Max output98K16K
Input $ / M tokens$0.20$2
Output $ / M tokens$1.10$6
Results tracked2751

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

GLM-4.5-Air: 33.3 (#259), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGLM-4.5-AirMistral Large
LMArena Coding13971277
SciCode—36.2%
GSO2.9%—
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

GLM-4.5-Air: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGLM-4.5-AirMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning GLM-4.5-Air leads

GLM-4.5-Air: 24.1 (#166), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGLM-4.5-AirMistral Large
LMArena Hard Prompts13791257
ForecastBench59.257.1
SimpleBench—22.5%
Kagi LLM Benchmark43%—
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
LiveBench—48.4%

Math GLM-4.5-Air leads

GLM-4.5-Air: 36.2 (#170), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGLM-4.5-AirMistral Large
Omni-MATH39.1%28.1%
LMArena Math13961262
OTIS Mock AIME 2024-2025—8.5%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge GLM-4.5-Air leads

GLM-4.5-Air: 35.0 (#191), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGLM-4.5-AirMistral Large
MMLU-Pro76.2%59.9%
Vectara Hallucination Rate9.3%4.5%
GPQA (HELM)59.4%43.5%
LMArena Expert13701232
GPQA Diamond—51.3%
Humanity's Last Exam8.1%—
Confabulations—21.4%
MMLU—80%

Multilingual GLM-4.5-Air leads

GLM-4.5-Air: 49.1 (#135), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGLM-4.5-AirMistral Large
LMArena Non-English13661237
LMArena Chinese14261240
LMArena French13991325
LMArena German13771254
LMArena Japanese13481188
LMArena Korean13081202
LMArena Russian13731257
LMArena Spanish13861268

Instruction Following GLM-4.5-Air leads

GLM-4.5-Air: 69.6 (#171), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGLM-4.5-AirMistral Large
IFEval81.2%87.7%
LMArena Instruction Following13541249
LiveBench Instruction Following—67.9%

Long Context GLM-4.5-Air leads

GLM-4.5-Air: 41.6 (#135), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGLM-4.5-AirMistral Large
LMArena Longer Query13661261

Writing & Preference GLM-4.5-Air leads

GLM-4.5-Air: 55.9 (#139), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGLM-4.5-AirMistral Large
LMArena Text13841266
LMArena Creative Writing13431243
WildBench78.9%80.1%
LMArena Multi-Turn13711260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
LiveBench Language—39.4%

Frequently asked questions

Is GLM-4.5-Air better than Mistral Large?

GLM-4.5-Air is the stronger model overall, scoring 38.9 to 31.9 on the Noometry Index.

Which is cheaper, GLM-4.5-Air or Mistral Large?

GLM-4.5-Air is cheaper. It lists at $0.20 per million input tokens and $1.10 per million output tokens; Mistral Large lists at $2 and $6.

Is GLM-4.5-Air or Mistral Large better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 33.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do GLM-4.5-Air and Mistral Large share?

24 benchmarks have published results for both models. GLM-4.5-Air has 27 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper