Model comparison

Gemini 1.5 Flash (May 2024) vs GLM-4.6

GLM-4.6 is the stronger model overall, scoring 41.4 to 33.2 on the Noometry Index.

Last verified . 18 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

GLM-4.6 Z.ai (Zhipu)

41.4

Rank #135 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 0 categories and GLM-4.6 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-4.6 leads 39.1 to 22.1.
  • GLM-4.6 has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash (May 2024) and GLM-4.6 specifications
Gemini 1.5 Flash (May 2024)GLM-4.6
ProviderGoogleZ.ai (Zhipu)
Noometry Index33.241.4
Released2024-05-142025-09-30
WeightsProprietaryOpen
Context window—205K
Max output—131K
Input $ / M tokens—$0.60
Output $ / M tokens—$2.20
Results tracked4229

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.6 leads

Gemini 1.5 Flash (May 2024): 34.4 (#236), GLM-4.6: 40.1 (#148)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GLM-4.6
LMArena Coding12611449
SWE-bench Verified (bash only)—55.4%
LMArena WebDev—1340
SciCode—38.4%
WeirdML24.9%—
BigCodeBench Instruct43.5%—
BigCodeBench Complete55.1%—
ALE-Bench—340.82
HumanEval+75.6%—
MBPP+67.5%—

Agentic & Tool Use GLM-4.6 leads

Gemini 1.5 Flash (May 2024): 26.6 (#102), GLM-4.6: 32.3 (#66)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GLM-4.6
Terminal-Bench—24.5%
Berkeley Function Calling Leaderboard—72.4%
BALROG14.6%—

Reasoning GLM-4.6 leads

Gemini 1.5 Flash (May 2024): 21.7 (#215), GLM-4.6: 23.7 (#172)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GLM-4.6
LMArena Hard Prompts12571440
Kagi LLM Benchmark—47.4%
CritPt—1.1%
DTBench53.8%—
Epoch Capabilities Index129.36—
ForecastBench53.9—
PIQA87.5%—

Math GLM-4.6 leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), GLM-4.6: 39.1 (#111)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GLM-4.6
LMArena Math12691432
FrontierMath (Feb 2025 set)0%3.8%
OTIS Mock AIME 2024-202516.3%—
Omni-MATH30.4%—
MATH Level 561.9%—
FrontierMath Tier 4 (v1)—2.1%
GSM8K82.4%—

Knowledge GLM-4.6 leads

Gemini 1.5 Flash (May 2024): 26.2 (#260), GLM-4.6: 40.2 (#124)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GLM-4.6
LMArena Expert12331431
GPQA Diamond47.3%—
MMLU-Pro67.8%—
Vectara Hallucination Rate—9.5%
GPQA (HELM)43.7%—
BoolQ85.8%—
MMLU77.9%—

Multimodal Not comparable

Gemini 1.5 Flash (May 2024): 36.0 (#81), GLM-4.6: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GLM-4.6
LMArena Vision1141—
Video-MME70.3%—
GeoBench76%—

Multilingual GLM-4.6 leads

Gemini 1.5 Flash (May 2024): 42.9 (#189), GLM-4.6: 53.5 (#66)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GLM-4.6
LMArena Non-English12781426
LMArena Chinese12951499
LMArena French12581459
LMArena German12621447
LMArena Japanese12521393
LMArena Korean12211400
LMArena Russian12881419
LMArena Spanish12431436

Instruction Following GLM-4.6 leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), GLM-4.6: 74.3 (#98)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GLM-4.6
LMArena Instruction Following12581410
IFEval83.1%—

Long Context GLM-4.6 leads

Gemini 1.5 Flash (May 2024): 39.0 (#187), GLM-4.6: 43.4 (#94)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GLM-4.6
LMArena Longer Query12841422

Writing & Preference GLM-4.6 leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), GLM-4.6: 61.1 (#90)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GLM-4.6
LMArena Text12871440
LMArena Creative Writing12851411
LMArena Multi-Turn12531427
EQ-Bench Creative Writing—1411
WildBench79.2%—

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than GLM-4.6?

GLM-4.6 is the stronger model overall, scoring 41.4 to 33.2 on the Noometry Index.

Is Gemini 1.5 Flash (May 2024) or GLM-4.6 better for coding?

GLM-4.6 scores higher on coding benchmarks: 40.1 versus 34.4 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash (May 2024) and GLM-4.6 share?

18 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and GLM-4.6 has 29.

Related comparisons

Go deeper