Model comparison

Gemini 1.5 Flash (May 2024) vs Qwen2.5 72B Instruct

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 31.9 on the Noometry Index.

Last verified . 34 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 34 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 7 categories and Qwen2.5 72B Instruct in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Gemini 1.5 Flash (May 2024) leads 26.6 to 22.1.
  • The biggest single-benchmark swing is DTBench: 53.8% for Gemini 1.5 Flash (May 2024) and 62.9% for Qwen2.5 72B Instruct.
  • Qwen2.5 72B Instruct has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash (May 2024) and Qwen2.5 72B Instruct specifications
Gemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
ProviderGoogleAlibaba (Qwen)
Noometry Index33.231.9
Released2024-05-142024-09
WeightsProprietaryOpen
Context window—131K
Max output—8K
Input $ / M tokens—$1.40
Output $ / M tokens—$5.60
Results tracked4243

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 34.4 (#236), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
WeirdML24.9%16%
BigCodeBench Instruct43.5%45.8%
LMArena Coding12611292
BigCodeBench Complete55.1%55.9%
HumanEval+75.6%—
MBPP+67.5%—

Agentic & Tool Use Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 26.6 (#102), Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
BALROG14.6%16.2%
TheAgentCompany—5.7%
METR Time Horizons—35.8%

Reasoning Too close to call

Gemini 1.5 Flash (May 2024): 21.7 (#215), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
LMArena Hard Prompts12571271
DTBench53.8%62.9%
Epoch Capabilities Index129.36129
ForecastBench53.957.5
PIQA87.5%82.6%
LMCA—13.4%
BIG-Bench Hard—79.8%
HellaSwag—84.8%
WinoGrande—82.3%

Math Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
OTIS Mock AIME 2024-202516.3%8.1%
Omni-MATH30.4%33%
LMArena Math12691283
MATH Level 561.9%63.2%
FrontierMath (Feb 2025 set)0%—
GSM8K82.4%—

Knowledge Too close to call

Gemini 1.5 Flash (May 2024): 26.2 (#260), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
GPQA Diamond47.3%49.1%
MMLU-Pro67.8%63.1%
GPQA (HELM)43.7%42.6%
LMArena Expert12331245
MMLU77.9%85.3%
Confabulations—19.1%
ARC (AI2) Challenge—94.5%
BoolQ85.8%—
TriviaQA—71.9%

Multimodal Not comparable

Gemini 1.5 Flash (May 2024): 36.0 (#81), Qwen2.5 72B Instruct: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
LMArena Vision1141—
Video-MME70.3%—
GeoBench76%—

Multilingual Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 42.9 (#189), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
LMArena Non-English12781252
LMArena Chinese12951272
LMArena French12581280
LMArena German12621234
LMArena Japanese12521180
LMArena Korean12211188
LMArena Russian12881264
LMArena Spanish12431256

Instruction Following Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
IFEval83.1%80.6%
LMArena Instruction Following12581254

Long Context Too close to call

Gemini 1.5 Flash (May 2024): 39.0 (#187), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
LMArena Longer Query12841282

Writing & Preference Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5 72B Instruct
LMArena Text12871269
LMArena Creative Writing12851221
WildBench79.2%80.2%
LMArena Multi-Turn12531272

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than Qwen2.5 72B Instruct?

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 31.9 on the Noometry Index.

Is Gemini 1.5 Flash (May 2024) or Qwen2.5 72B Instruct better for coding?

Gemini 1.5 Flash (May 2024) scores higher on coding benchmarks: 34.4 versus 33.2 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash (May 2024) and Qwen2.5 72B Instruct share?

34 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper