Model comparison

GLM-4.6 vs Nova 2 Lite

GLM-4.6 is the stronger model overall, scoring 41.4 to 39.7 on the Noometry Index.

Last verified . 19 shared benchmarks.

GLM-4.6 Z.ai (Zhipu)

41.4

Rank #135 Confirmed

Nova 2 Lite Amazon

39.7

Rank #161 Confirmed

Summary

  • They share 19 benchmarks with published results for both. GLM-4.6 scores higher in 6 categories and Nova 2 Lite in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where GLM-4.6 leads 32.3 to 24.1.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 72.4% for GLM-4.6 and 27.1% for Nova 2 Lite.
  • Nova 2 Lite is cheaper at $0.30 / $2.50 per million input/output tokens, against $0.60 / $2.20 for GLM-4.6.
  • Nova 2 Lite accepts more context: 1M tokens versus 205K.
  • GLM-4.6 has downloadable open weights; the other is API-only.

Side by side

GLM-4.6 and Nova 2 Lite specifications
GLM-4.6Nova 2 Lite
ProviderZ.ai (Zhipu)Amazon
Noometry Index41.439.7
Released2025-09-302025-12-01
WeightsOpenProprietary
Context window205K1M
Max output131K64K
Input $ / M tokens$0.60$0.30
Output $ / M tokens$2.20$2.50
Results tracked2919

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GLM-4.6: 40.1 (#148), Nova 2 Lite: 40.7 (#134)

Coding benchmarks
BenchmarkGLM-4.6Nova 2 Lite
LMArena Coding14491385
SWE-bench Verified (bash only)55.4%—
LMArena WebDev1340—
SciCode38.4%—
ALE-Bench340.82—

Agentic & Tool Use GLM-4.6 leads

GLM-4.6: 32.3 (#66), Nova 2 Lite: 24.1 (#121)

Agentic & Tool Use benchmarks
BenchmarkGLM-4.6Nova 2 Lite
Berkeley Function Calling Leaderboard72.4%27.1%
Terminal-Bench24.5%—

Reasoning Nova 2 Lite leads

GLM-4.6: 23.7 (#172), Nova 2 Lite: 27.5 (#118)

Reasoning benchmarks
BenchmarkGLM-4.6Nova 2 Lite
LMArena Hard Prompts14401364
Kagi LLM Benchmark47.4%—
CritPt1.1%—

Math GLM-4.6 leads

GLM-4.6: 39.1 (#111), Nova 2 Lite: 37.5 (#156)

Math benchmarks
BenchmarkGLM-4.6Nova 2 Lite
LMArena Math14321359
FrontierMath (Feb 2025 set)3.8%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge Nova 2 Lite leads

GLM-4.6: 40.2 (#124), Nova 2 Lite: 43.0 (#94)

Knowledge benchmarks
BenchmarkGLM-4.6Nova 2 Lite
Vectara Hallucination Rate9.5%5.1%
LMArena Expert14311358

Multilingual GLM-4.6 leads

GLM-4.6: 53.5 (#66), Nova 2 Lite: 47.1 (#153)

Multilingual benchmarks
BenchmarkGLM-4.6Nova 2 Lite
LMArena Non-English14261337
LMArena Chinese14991364
LMArena French14591381
LMArena German14471343
LMArena Japanese13931271
LMArena Korean14001284
LMArena Russian14191343
LMArena Spanish14361373

Instruction Following GLM-4.6 leads

GLM-4.6: 74.3 (#98), Nova 2 Lite: 70.5 (#161)

Instruction Following benchmarks
BenchmarkGLM-4.6Nova 2 Lite
LMArena Instruction Following14101335

Long Context GLM-4.6 leads

GLM-4.6: 43.4 (#94), Nova 2 Lite: 40.6 (#150)

Long Context benchmarks
BenchmarkGLM-4.6Nova 2 Lite
LMArena Longer Query14221335

Writing & Preference GLM-4.6 leads

GLM-4.6: 61.1 (#90), Nova 2 Lite: 53.9 (#154)

Writing & Preference benchmarks
BenchmarkGLM-4.6Nova 2 Lite
LMArena Text14401362
LMArena Creative Writing14111291
LMArena Multi-Turn14271338
EQ-Bench Creative Writing1411—

Frequently asked questions

Is GLM-4.6 better than Nova 2 Lite?

GLM-4.6 is the stronger model overall, scoring 41.4 to 39.7 on the Noometry Index.

Which is cheaper, GLM-4.6 or Nova 2 Lite?

Nova 2 Lite is cheaper. It lists at $0.30 per million input tokens and $2.50 per million output tokens; GLM-4.6 lists at $0.60 and $2.20.

Is GLM-4.6 or Nova 2 Lite better for coding?

They score almost the same on coding (40.1 vs 40.7); test both on your own repository before choosing.

Which has the bigger context window?

Nova 2 Lite does, with 1M tokens against 205K.

How many benchmarks do GLM-4.6 and Nova 2 Lite share?

19 benchmarks have published results for both models. GLM-4.6 has 29 scored results on Noometry and Nova 2 Lite has 19.

Related comparisons

Go deeper