Model comparison

Amazon Nova Experimental Chat 12 10 vs GLM-4.7

Amazon Nova Experimental Chat 12 10 and GLM-4.7 score almost the same on the Noometry Index (42.9 vs 42.0), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

GLM-4.7 Z.ai (Zhipu)

42.0

Rank #124 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Amazon Nova Experimental Chat 12 10 scores higher in 2 categories and GLM-4.7 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GLM-4.7 leads 47.0 to 39.2.
  • GLM-4.7 has downloadable open weights; the other is API-only.

Side by side

Amazon Nova Experimental Chat 12 10 and GLM-4.7 specifications
Amazon Nova Experimental Chat 12 10GLM-4.7
ProviderAmazonZ.ai (Zhipu)
Noometry Index42.942.0
Released—2025-12-22
WeightsProprietaryOpen
Context window—205K
Max output—131K
Input $ / M tokens—$0.60
Output $ / M tokens—$2.20
Results tracked1236

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.7 leads

Amazon Nova Experimental Chat 12 10: 42.2 (#110), GLM-4.7: 44.0 (#79)

Coding benchmarks
BenchmarkAmazon Nova Experimental Chat 12 10GLM-4.7
LMArena Coding14301454
LMArena WebDev—1435
SciCode—45.1%
ALE-Bench—399.48

Agentic & Tool Use Not comparable

Amazon Nova Experimental Chat 12 10: —, GLM-4.7: 26.5 (#103)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova Experimental Chat 12 10GLM-4.7
Terminal-Bench—33.4%
Vending-Bench 2—2,377

Reasoning Amazon Nova Experimental Chat 12 10 leads

Amazon Nova Experimental Chat 12 10: 29.3 (#94), GLM-4.7: 24.3 (#164)

Reasoning benchmarks
BenchmarkAmazon Nova Experimental Chat 12 10GLM-4.7
LMArena Hard Prompts14271443
SimpleBench—47.7%
CritPt—1.7%
Chess Puzzles—6%
Epoch Capabilities Index—143.51

Math Too close to call

Amazon Nova Experimental Chat 12 10: 38.9 (#124), GLM-4.7: 38.6 (#135)

Math benchmarks
BenchmarkAmazon Nova Experimental Chat 12 10GLM-4.7
LMArena Math14191423
OTIS Mock AIME 2024-2025—83.3%
ProofBench—6%
FrontierMath (Feb 2025 set)—2.4%
FrontierMath Tier 4 (v1)—0%

Knowledge GLM-4.7 leads

Amazon Nova Experimental Chat 12 10: 39.2 (#136), GLM-4.7: 47.0 (#80)

Knowledge benchmarks
BenchmarkAmazon Nova Experimental Chat 12 10GLM-4.7
LMArena Expert14081424
GPQA Diamond—83.3%
SimpleQA Verified—32.2%
Vectara Hallucination Rate—11.7%

Multilingual GLM-4.7 leads

Amazon Nova Experimental Chat 12 10: 51.4 (#108), GLM-4.7: 52.8 (#79)

Multilingual benchmarks
BenchmarkAmazon Nova Experimental Chat 12 10GLM-4.7
LMArena Non-English13981417
LMArena Chinese14401495
LMArena Russian13951423
LMArena French—1432
LMArena German—1424
LMArena Japanese—1439
LMArena Korean—1399
LMArena Spanish—1434

Instruction Following GLM-4.7 leads

Amazon Nova Experimental Chat 12 10: 73.2 (#123), GLM-4.7: 74.4 (#95)

Instruction Following benchmarks
BenchmarkAmazon Nova Experimental Chat 12 10GLM-4.7
LMArena Instruction Following13871411

Long Context Too close to call

Amazon Nova Experimental Chat 12 10: 42.4 (#125), GLM-4.7: 42.8 (#116)

Long Context benchmarks
BenchmarkAmazon Nova Experimental Chat 12 10GLM-4.7
LMArena Longer Query13911432
CL-bench—15.9%
CL-bench Life—10.9%

Writing & Preference GLM-4.7 leads

Amazon Nova Experimental Chat 12 10: 59.5 (#110), GLM-4.7: 60.9 (#93)

Writing & Preference benchmarks
BenchmarkAmazon Nova Experimental Chat 12 10GLM-4.7
LMArena Text14201435
LMArena Creative Writing13501401
LMArena Multi-Turn14091446
EQ-Bench Creative Writing—1413

Frequently asked questions

Is Amazon Nova Experimental Chat 12 10 better than GLM-4.7?

Amazon Nova Experimental Chat 12 10 and GLM-4.7 score almost the same on the Noometry Index (42.9 vs 42.0), so choose on price, context window or the category you care about most.

Is Amazon Nova Experimental Chat 12 10 or GLM-4.7 better for coding?

GLM-4.7 scores higher on coding benchmarks: 44.0 versus 42.2 in the Noometry coding category.

How many benchmarks do Amazon Nova Experimental Chat 12 10 and GLM-4.7 share?

12 benchmarks have published results for both models. Amazon Nova Experimental Chat 12 10 has 12 scored results on Noometry and GLM-4.7 has 36.

Related comparisons

Go deeper