Model comparison

GLM-4.5-Air vs Llama 3-70B

GLM-4.5-Air is the stronger model overall, scoring 38.9 to 28.8 on the Noometry Index.

Last verified . 19 shared benchmarks.

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

Llama 3-70B Meta

28.8

Rank #323 Confirmed

Summary

  • They share 19 benchmarks with published results for both. GLM-4.5-Air scores higher in 7 categories and Llama 3-70B in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-4.5-Air leads 36.2 to 12.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 43% for GLM-4.5-Air and 35.1% for Llama 3-70B.

Side by side

GLM-4.5-Air and Llama 3-70B specifications
GLM-4.5-AirLlama 3-70B
ProviderZ.ai (Zhipu)Meta
Noometry Index38.928.8
Released2025-07-202024-04-18
WeightsOpenOpen
Context window131K—
Max output98K—
Input $ / M tokens$0.20—
Output $ / M tokens$1.10—
Results tracked2731

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3-70B leads

GLM-4.5-Air: 33.3 (#259), Llama 3-70B: 35.8 (#218)

Coding benchmarks
BenchmarkGLM-4.5-AirLlama 3-70B
LMArena Coding13971206
GSO2.9%—
BigCodeBench Instruct—43.6%
BigCodeBench Complete—54.5%
HumanEval+—72%
MBPP+—69%

Agentic & Tool Use Not comparable

GLM-4.5-Air: —, Llama 3-70B: 21.1 (#139)

Agentic & Tool Use benchmarks
BenchmarkGLM-4.5-AirLlama 3-70B
Cybench—5%

Reasoning GLM-4.5-Air leads

GLM-4.5-Air: 24.1 (#166), Llama 3-70B: 18.0 (#288)

Reasoning benchmarks
BenchmarkGLM-4.5-AirLlama 3-70B
Kagi LLM Benchmark43%35.1%
LMArena Hard Prompts13791195
ForecastBench59.257.1
DTBench—54.2%
Epoch Capabilities Index—122.93
WinoGrande—83.5%

Math GLM-4.5-Air leads

GLM-4.5-Air: 36.2 (#170), Llama 3-70B: 12.8 (#305)

Math benchmarks
BenchmarkGLM-4.5-AirLlama 3-70B
LMArena Math13961218
OTIS Mock AIME 2024-2025—4.3%
Omni-MATH39.1%—
MATH Level 5—22.6%

Knowledge GLM-4.5-Air leads

GLM-4.5-Air: 35.0 (#191), Llama 3-70B: 20.8 (#277)

Knowledge benchmarks
BenchmarkGLM-4.5-AirLlama 3-70B
LMArena Expert13701149
GPQA Diamond—40.6%
Humanity's Last Exam8.1%—
MMLU-Pro76.2%—
Vectara Hallucination Rate9.3%—
GPQA (HELM)59.4%—
MMLU—79.3%

Multilingual GLM-4.5-Air leads

GLM-4.5-Air: 49.1 (#135), Llama 3-70B: 33.6 (#251)

Multilingual benchmarks
BenchmarkGLM-4.5-AirLlama 3-70B
LMArena Non-English13661142
LMArena Chinese14261114
LMArena French13991232
LMArena German13771169
LMArena Japanese13481017
LMArena Korean13081017
LMArena Russian13731159
LMArena Spanish13861241

Instruction Following GLM-4.5-Air leads

GLM-4.5-Air: 69.6 (#171), Llama 3-70B: 62.5 (#238)

Instruction Following benchmarks
BenchmarkGLM-4.5-AirLlama 3-70B
LMArena Instruction Following13541194
IFEval81.2%—

Long Context GLM-4.5-Air leads

GLM-4.5-Air: 41.6 (#135), Llama 3-70B: 35.6 (#240)

Long Context benchmarks
BenchmarkGLM-4.5-AirLlama 3-70B
LMArena Longer Query13661174

Writing & Preference GLM-4.5-Air leads

GLM-4.5-Air: 55.9 (#139), Llama 3-70B: 42.8 (#231)

Writing & Preference benchmarks
BenchmarkGLM-4.5-AirLlama 3-70B
LMArena Text13841221
LMArena Creative Writing13431210
LMArena Multi-Turn13711223
WildBench78.9%—

Frequently asked questions

Is GLM-4.5-Air better than Llama 3-70B?

GLM-4.5-Air is the stronger model overall, scoring 38.9 to 28.8 on the Noometry Index.

Is GLM-4.5-Air or Llama 3-70B better for coding?

Llama 3-70B scores higher on coding benchmarks: 35.8 versus 33.3 in the Noometry coding category.

How many benchmarks do GLM-4.5-Air and Llama 3-70B share?

19 benchmarks have published results for both models. GLM-4.5-Air has 27 scored results on Noometry and Llama 3-70B has 31.

Related comparisons

Go deeper