Model comparison

GLM-4.5-Air vs Nvidia Llama 3.3 Nemotron Super 49b v1.5

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 38.9 on the Noometry Index.

Last verified . 12 shared benchmarks.

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

Summary

  • They share 12 benchmarks with published results for both. GLM-4.5-Air scores higher in 4 categories and Nvidia Llama 3.3 Nemotron Super 49b v1.5 in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads 39.8 to 33.3.
  • Nvidia Llama 3.3 Nemotron Super 49b v1.5 is cheaper at $0.40 / $0.40 per million input/output tokens, against $0.20 / $1.10 for GLM-4.5-Air.

Side by side

GLM-4.5-Air and Nvidia Llama 3.3 Nemotron Super 49b v1.5 specifications
GLM-4.5-AirNvidia Llama 3.3 Nemotron Super 49b v1.5
ProviderZ.ai (Zhipu)NVIDIA
Noometry Index38.940.3
Released2025-07-202025-07-25
WeightsOpenOpen
Context window131K131K
Max output98K131K
Input $ / M tokens$0.20$0.40
Output $ / M tokens$1.10$0.40
Results tracked2712

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

GLM-4.5-Air: 33.3 (#259), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 39.8 (#154)

Coding benchmarks
BenchmarkGLM-4.5-AirNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Coding13971355
GSO2.9%—

Reasoning Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

GLM-4.5-Air: 24.1 (#166), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 26.8 (#128)

Reasoning benchmarks
BenchmarkGLM-4.5-AirNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Hard Prompts13791336
Kagi LLM Benchmark43%—
ForecastBench59.2—

Math Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

GLM-4.5-Air: 36.2 (#170), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 38.2 (#141)

Math benchmarks
BenchmarkGLM-4.5-AirNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Math13961392
Omni-MATH39.1%—

Knowledge Nvidia Llama 3.3 Nemotron Super 49b v1.5 leads

GLM-4.5-Air: 35.0 (#191), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 36.7 (#165)

Knowledge benchmarks
BenchmarkGLM-4.5-AirNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Expert13701330
Humanity's Last Exam8.1%—
MMLU-Pro76.2%—
Vectara Hallucination Rate9.3%—
GPQA (HELM)59.4%—

Multilingual GLM-4.5-Air leads

GLM-4.5-Air: 49.1 (#135), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 45.5 (#168)

Multilingual benchmarks
BenchmarkGLM-4.5-AirNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Non-English13661316
LMArena Japanese13481300
LMArena Russian13731332
LMArena Chinese1426—
LMArena French1399—
LMArena German1377—
LMArena Korean1308—
LMArena Spanish1386—

Instruction Following GLM-4.5-Air leads

GLM-4.5-Air: 69.6 (#171), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 68.6 (#188)

Instruction Following benchmarks
BenchmarkGLM-4.5-AirNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Instruction Following13541299
IFEval81.2%—

Long Context GLM-4.5-Air leads

GLM-4.5-Air: 41.6 (#135), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 40.0 (#164)

Long Context benchmarks
BenchmarkGLM-4.5-AirNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Longer Query13661315

Writing & Preference GLM-4.5-Air leads

GLM-4.5-Air: 55.9 (#139), Nvidia Llama 3.3 Nemotron Super 49b v1.5: 53.1 (#159)

Writing & Preference benchmarks
BenchmarkGLM-4.5-AirNvidia Llama 3.3 Nemotron Super 49b v1.5
LMArena Text13841338
LMArena Creative Writing13431307
LMArena Multi-Turn13711334
WildBench78.9%—

Frequently asked questions

Is GLM-4.5-Air better than Nvidia Llama 3.3 Nemotron Super 49b v1.5?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is the stronger model overall, scoring 40.3 to 38.9 on the Noometry Index.

Which is cheaper, GLM-4.5-Air or Nvidia Llama 3.3 Nemotron Super 49b v1.5?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 is cheaper. It lists at $0.40 per million input tokens and $0.40 per million output tokens; GLM-4.5-Air lists at $0.20 and $1.10.

Is GLM-4.5-Air or Nvidia Llama 3.3 Nemotron Super 49b v1.5 better for coding?

Nvidia Llama 3.3 Nemotron Super 49b v1.5 scores higher on coding benchmarks: 39.8 versus 33.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do GLM-4.5-Air and Nvidia Llama 3.3 Nemotron Super 49b v1.5 share?

12 benchmarks have published results for both models. GLM-4.5-Air has 27 scored results on Noometry and Nvidia Llama 3.3 Nemotron Super 49b v1.5 has 12.

Related comparisons

Go deeper