Model comparison

Gemini 1.5 Flash (May 2024) vs Qwen2.5-Coder-32B

Gemini 1.5 Flash (May 2024) and Qwen2.5-Coder-32B score almost the same on the Noometry Index (33.2 vs 33.4), so choose on price, context window or the category you care about most.

Last verified . 19 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 6 categories and Qwen2.5-Coder-32B in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Gemini 1.5 Flash (May 2024) leads 34.4 to 22.6.
  • The biggest single-benchmark swing is BigCodeBench Instruct: 43.5% for Gemini 1.5 Flash (May 2024) and 49% for Qwen2.5-Coder-32B.
  • Qwen2.5-Coder-32B has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash (May 2024) and Qwen2.5-Coder-32B specifications
Gemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
ProviderGoogleAlibaba (Qwen)
Noometry Index33.233.4
Released2024-05-142024-09-18
WeightsProprietaryOpen
Context window—33K
Max output—29K
Input $ / M tokens—$0.66
Output $ / M tokens—$1
Results tracked4231

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 34.4 (#236), Qwen2.5-Coder-32B: 22.6 (#333)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
BigCodeBench Instruct43.5%49%
LMArena Coding12611276
BigCodeBench Complete55.1%58%
HumanEval+75.6%87.2%
MBPP+67.5%77%
SWE-bench Verified (bash only)—9%
Aider Polyglot—16.4%
WeirdML24.9%—
LiveBench Coding—56.9%

Agentic & Tool Use Not comparable

Gemini 1.5 Flash (May 2024): 26.6 (#102), Qwen2.5-Coder-32B: —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
BALROG14.6%—

Reasoning Too close to call

Gemini 1.5 Flash (May 2024): 21.7 (#215), Qwen2.5-Coder-32B: 21.2 (#225)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
LMArena Hard Prompts12571251
Epoch Capabilities Index129.36119.49
LiveBench Reasoning—42.1%
DTBench53.8%—
LiveBench Data Analysis—49.9%
ForecastBench53.9—
HellaSwag—83%
LiveBench—46.2%
PIQA87.5%—
WinoGrande—80.8%

Math Qwen2.5-Coder-32B leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), Qwen2.5-Coder-32B: 33.3 (#204)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
LMArena Math12691251
GSM8K82.4%93%
OTIS Mock AIME 2024-202516.3%—
Omni-MATH30.4%—
LiveBench Math—46.6%
MATH Level 561.9%—
FrontierMath (Feb 2025 set)0%—

Knowledge Qwen2.5-Coder-32B leads

Gemini 1.5 Flash (May 2024): 26.2 (#260), Qwen2.5-Coder-32B: 33.4 (#203)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
LMArena Expert12331221
MMLU77.9%79.1%
GPQA Diamond47.3%—
MMLU-Pro67.8%—
GPQA (HELM)43.7%—
ARC (AI2) Challenge—70.5%
BoolQ85.8%—

Multimodal Not comparable

Gemini 1.5 Flash (May 2024): 36.0 (#81), Qwen2.5-Coder-32B: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
LMArena Vision1141—
Video-MME70.3%—
GeoBench76%—

Multilingual Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 42.9 (#189), Qwen2.5-Coder-32B: 37.8 (#235)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
LMArena Non-English12781205
LMArena Chinese12951222
LMArena Russian12881228
LMArena French1258—
LMArena German1262—
LMArena Japanese1252—
LMArena Korean1221—
LMArena Spanish1243—

Instruction Following Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), Qwen2.5-Coder-32B: 61.4 (#245)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
LMArena Instruction Following12581223
LiveBench Instruction Following—58.7%
IFEval83.1%—

Long Context Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 39.0 (#187), Qwen2.5-Coder-32B: 38.0 (#208)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
LMArena Longer Query12841251

Writing & Preference Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), Qwen2.5-Coder-32B: 41.6 (#240)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Qwen2.5-Coder-32B
LMArena Text12871230
LMArena Creative Writing12851174
LMArena Multi-Turn12531222
WildBench79.2%—
LiveBench Language—23.3%

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than Qwen2.5-Coder-32B?

Gemini 1.5 Flash (May 2024) and Qwen2.5-Coder-32B score almost the same on the Noometry Index (33.2 vs 33.4), so choose on price, context window or the category you care about most.

Is Gemini 1.5 Flash (May 2024) or Qwen2.5-Coder-32B better for coding?

Gemini 1.5 Flash (May 2024) scores higher on coding benchmarks: 34.4 versus 22.6 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash (May 2024) and Qwen2.5-Coder-32B share?

19 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and Qwen2.5-Coder-32B has 31.

Related comparisons

Go deeper