Model comparison

Gemini 2.5 Flash-Lite vs GLM-4.7

GLM-4.7 is the stronger model overall, scoring 42.0 to 37.0 on the Noometry Index. Gemini 2.5 Flash-Lite costs 5.7× less per token, which makes it the better buy when GLM-4.7's lead doesn't matter for your workload.

Last verified . 20 shared benchmarks.

Gemini 2.5 Flash-Lite Google

37.0

Rank #211 Confirmed

GLM-4.7 Z.ai (Zhipu)

42.0

Rank #124 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Gemini 2.5 Flash-Lite scores higher in 1 category and GLM-4.7 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GLM-4.7 leads 47.0 to 32.5.
  • The biggest single-benchmark swing is Vectara Hallucination Rate: 3.3% for Gemini 2.5 Flash-Lite and 11.7% for GLM-4.7.
  • Gemini 2.5 Flash-Lite is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.60 / $2.20 for GLM-4.7.
  • Gemini 2.5 Flash-Lite accepts more context: 1.05M tokens versus 205K.
  • GLM-4.7 has downloadable open weights; the other is API-only.

Side by side

Gemini 2.5 Flash-Lite and GLM-4.7 specifications
Gemini 2.5 Flash-LiteGLM-4.7
ProviderGoogleZ.ai (Zhipu)
Noometry Index37.042.0
Released2025-06-172025-12-22
WeightsProprietaryOpen
Context window1.05M205K
Max output66K131K
Input $ / M tokens$0.10$0.60
Output $ / M tokens$0.40$2.20
Results tracked3336

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.7 leads

Gemini 2.5 Flash-Lite: 38.5 (#173), GLM-4.7: 44.0 (#79)

Coding benchmarks
BenchmarkGemini 2.5 Flash-LiteGLM-4.7
LMArena Coding13731454
ALE-Bench325.9399.48
LMArena WebDev—1435
SciCode—45.1%
WeirdML35.2%—

Agentic & Tool Use Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 28.0 (#96), GLM-4.7: 26.5 (#103)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 Flash-LiteGLM-4.7
Terminal-Bench—33.4%
Berkeley Function Calling Leaderboard36.9%—
Vending-Bench 2—2,377

Reasoning GLM-4.7 leads

Gemini 2.5 Flash-Lite: 22.2 (#205), GLM-4.7: 24.3 (#164)

Reasoning benchmarks
BenchmarkGemini 2.5 Flash-LiteGLM-4.7
LMArena Hard Prompts13771443
Epoch Capabilities Index133.94143.51
SimpleBench—47.7%
Kagi LLM Benchmark40.5%—
CritPt—1.7%
Chess Puzzles—6%
DTBench62.8%—
LMCA18.1%—

Math Too close to call

Gemini 2.5 Flash-Lite: 38.0 (#144), GLM-4.7: 38.6 (#135)

Math benchmarks
BenchmarkGemini 2.5 Flash-LiteGLM-4.7
LMArena Math13731423
OTIS Mock AIME 2024-2025—83.3%
ProofBench—6%
Omni-MATH48%—
FrontierMath (Feb 2025 set)—2.4%
FrontierMath Tier 4 (v1)—0%

Knowledge GLM-4.7 leads

Gemini 2.5 Flash-Lite: 32.5 (#210), GLM-4.7: 47.0 (#80)

Knowledge benchmarks
BenchmarkGemini 2.5 Flash-LiteGLM-4.7
Vectara Hallucination Rate3.3%11.7%
LMArena Expert13731424
GPQA Diamond—83.3%
SimpleQA Verified—32.2%
MMLU-Pro53.7%—
GPQA (HELM)30.9%—

Multimodal Not comparable

Gemini 2.5 Flash-Lite: 29.1 (#114), GLM-4.7: —

Multimodal benchmarks
BenchmarkGemini 2.5 Flash-LiteGLM-4.7
LMArena Vision1198—
VPCT30%—

Multilingual GLM-4.7 leads

Gemini 2.5 Flash-Lite: 49.3 (#134), GLM-4.7: 52.8 (#79)

Multilingual benchmarks
BenchmarkGemini 2.5 Flash-LiteGLM-4.7
LMArena Non-English13691417
LMArena Chinese14041495
LMArena French13881432
LMArena German13891424
LMArena Japanese13591439
LMArena Korean13601399
LMArena Russian13731423
LMArena Spanish13961434

Instruction Following GLM-4.7 leads

Gemini 2.5 Flash-Lite: 70.0 (#168), GLM-4.7: 74.4 (#95)

Instruction Following benchmarks
BenchmarkGemini 2.5 Flash-LiteGLM-4.7
LMArena Instruction Following13671411
IFEval81%—

Long Context GLM-4.7 leads

Gemini 2.5 Flash-Lite: 33.3 (#262), GLM-4.7: 42.8 (#116)

Long Context benchmarks
BenchmarkGemini 2.5 Flash-LiteGLM-4.7
LMArena Longer Query13731432
Fiction.LiveBench47.2%—
CL-bench—15.9%
CL-bench Life—10.9%

Writing & Preference GLM-4.7 leads

Gemini 2.5 Flash-Lite: 56.8 (#135), GLM-4.7: 60.9 (#93)

Writing & Preference benchmarks
BenchmarkGemini 2.5 Flash-LiteGLM-4.7
LMArena Text13791435
LMArena Creative Writing13671401
LMArena Multi-Turn13661446
EQ-Bench Creative Writing—1413
WildBench81.8%—

Frequently asked questions

Is Gemini 2.5 Flash-Lite better than GLM-4.7?

GLM-4.7 is the stronger model overall, scoring 42.0 to 37.0 on the Noometry Index. Gemini 2.5 Flash-Lite costs 5.7× less per token, which makes it the better buy when GLM-4.7's lead doesn't matter for your workload.

Which is cheaper, Gemini 2.5 Flash-Lite or GLM-4.7?

Gemini 2.5 Flash-Lite is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; GLM-4.7 lists at $0.60 and $2.20.

Is Gemini 2.5 Flash-Lite or GLM-4.7 better for coding?

GLM-4.7 scores higher on coding benchmarks: 44.0 versus 38.5 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash-Lite does, with 1.05M tokens against 205K.

How many benchmarks do Gemini 2.5 Flash-Lite and GLM-4.7 share?

20 benchmarks have published results for both models. Gemini 2.5 Flash-Lite has 33 scored results on Noometry and GLM-4.7 has 36.

Related comparisons

Go deeper