Model comparison

Gemini 1.5 Pro (May 2024) vs QwQ-32B

QwQ-32B is the stronger model overall, scoring 39.8 to 32.1 on the Noometry Index.

Last verified . 24 shared benchmarks.

Gemini 1.5 Pro (May 2024) Google

32.1

Rank #261 Confirmed

QwQ-32B Alibaba (Qwen)

39.8

Rank #159 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Gemini 1.5 Pro (May 2024) scores higher in 2 categories and QwQ-32B in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where QwQ-32B leads 38.0 to 25.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 23.1% for Gemini 1.5 Pro (May 2024) and 59.2% for QwQ-32B.
  • QwQ-32B has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Pro (May 2024) and QwQ-32B specifications
Gemini 1.5 Pro (May 2024)QwQ-32B
ProviderGoogleAlibaba (Qwen)
Noometry Index32.139.8
Released2024-02-152024-11-28
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4536

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding QwQ-32B leads

Gemini 1.5 Pro (May 2024): 34.2 (#241), QwQ-32B: 35.4 (#226)

Coding benchmarks
BenchmarkGemini 1.5 Pro (May 2024)QwQ-32B
BigCodeBench Instruct43.8%44.6%
LMArena Coding12941333
BigCodeBench Complete57.5%54.4%
Aider Polyglot—20.9%
WeirdML22.2%—
LiveBench Coding—72.2%
CadEval34%—
HumanEval+79.3%—
MBPP+74.6%—

Agentic & Tool Use Not comparable

Gemini 1.5 Pro (May 2024): 17.9 (#145), QwQ-32B: —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Pro (May 2024)QwQ-32B
TheAgentCompany3.4%—
Cybench7.5%—
BALROG21%—

Reasoning QwQ-32B leads

Gemini 1.5 Pro (May 2024): 12.3 (#338), QwQ-32B: 23.7 (#174)

Reasoning benchmarks
BenchmarkGemini 1.5 Pro (May 2024)QwQ-32B
LMArena Hard Prompts12961325
Epoch Capabilities Index131.73137.6
ForecastBench58.458.3
ARC-AGI-20.8%—
SimpleBench27.1%—
Chess Puzzles—5%
LiveBench Reasoning—83.5%
DTBench59%—
LiveBench Data Analysis—65%
BIG-Bench Hard89.2%—
LiveBench—72%

Math QwQ-32B leads

Gemini 1.5 Pro (May 2024): 25.8 (#266), QwQ-32B: 38.0 (#143)

Math benchmarks
BenchmarkGemini 1.5 Pro (May 2024)QwQ-32B
OTIS Mock AIME 2024-202523.1%59.2%
LMArena Math13151359
Omni-MATH36.4%—
LiveBench Math—77.8%
MATH Level 570.4%—

Knowledge QwQ-32B leads

Gemini 1.5 Pro (May 2024): 29.4 (#239), QwQ-32B: 37.2 (#158)

Knowledge benchmarks
BenchmarkGemini 1.5 Pro (May 2024)QwQ-32B
GPQA Diamond57.2%65.3%
Confabulations13.5%15.6%
LMArena Expert12791324
Humanity's Last Exam4.6%—
MMLU-Pro73.7%—
GPQA (HELM)53.4%—
MMLU86.9%—

Multimodal Not comparable

Gemini 1.5 Pro (May 2024): 36.8 (#77), QwQ-32B: —

Multimodal benchmarks
BenchmarkGemini 1.5 Pro (May 2024)QwQ-32B
LMArena Vision1161—
Video-MME75%—

Multilingual Too close to call

Gemini 1.5 Pro (May 2024): 45.3 (#174), QwQ-32B: 44.8 (#176)

Multilingual benchmarks
BenchmarkGemini 1.5 Pro (May 2024)QwQ-32B
LMArena Non-English13121305
LMArena Chinese13311378
LMArena French13021336
LMArena German12861313
LMArena Japanese12921262
LMArena Korean12981279
LMArena Russian13201297
LMArena Spanish13111354

Instruction Following QwQ-32B leads

Gemini 1.5 Pro (May 2024): 68.6 (#185), QwQ-32B: 72.6 (#137)

Instruction Following benchmarks
BenchmarkGemini 1.5 Pro (May 2024)QwQ-32B
LMArena Instruction Following12971297
LiveBench Instruction Following—81.8%
IFEval83.7%—

Long Context QwQ-32B leads

Gemini 1.5 Pro (May 2024): 39.8 (#169), QwQ-32B: 49.0 (#11)

Long Context benchmarks
BenchmarkGemini 1.5 Pro (May 2024)QwQ-32B
LMArena Longer Query13081308
Fiction.LiveBench—83.3%

Writing & Preference Gemini 1.5 Pro (May 2024) leads

Gemini 1.5 Pro (May 2024): 52.4 (#172), QwQ-32B: 50.6 (#180)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Pro (May 2024)QwQ-32B
LMArena Text13191329
LMArena Creative Writing13331288
LMArena Multi-Turn12961314
Short-Story Creative Writing—80.2%
EQ-Bench Creative Writing—1257
WildBench81.3%—
LiveBench Language—51.4%

Frequently asked questions

Is Gemini 1.5 Pro (May 2024) better than QwQ-32B?

QwQ-32B is the stronger model overall, scoring 39.8 to 32.1 on the Noometry Index.

Is Gemini 1.5 Pro (May 2024) or QwQ-32B better for coding?

QwQ-32B scores higher on coding benchmarks: 35.4 versus 34.2 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Pro (May 2024) and QwQ-32B share?

24 benchmarks have published results for both models. Gemini 1.5 Pro (May 2024) has 45 scored results on Noometry and QwQ-32B has 36.

Related comparisons

Go deeper