Model comparison

Gemini 1.0 Pro vs Qwen1.5 4b Chat

Qwen1.5 4b Chat is the stronger model overall, scoring 28.8 to 27.3 on the Noometry Index.

Last verified . 13 shared benchmarks.

Gemini 1.0 Pro Google

27.3

Rank #332 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Gemini 1.0 Pro scores higher in 5 categories and Qwen1.5 4b Chat in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen1.5 4b Chat leads 30.4 to 9.3.
  • Qwen1.5 4b Chat has downloadable open weights; the other is API-only.

Side by side

Gemini 1.0 Pro and Qwen1.5 4b Chat specifications
Gemini 1.0 ProQwen1.5 4b Chat
ProviderGoogleAlibaba (Qwen)
Noometry Index27.328.8
Released2023-12-13—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2413

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.0 Pro leads

Gemini 1.0 Pro: 32.2 (#275), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkGemini 1.0 ProQwen1.5 4b Chat
LMArena Coding1108999
HumanEval+55.5%—
MBPP+61.4%—

Reasoning Qwen1.5 4b Chat leads

Gemini 1.0 Pro: 17.1 (#296), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkGemini 1.0 ProQwen1.5 4b Chat
LMArena Hard Prompts1109976
DTBench45.9%—
Epoch Capabilities Index117.04—

Math Qwen1.5 4b Chat leads

Gemini 1.0 Pro: 9.3 (#321), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkGemini 1.0 ProQwen1.5 4b Chat
LMArena Math11321026
OTIS Mock AIME 2024-20251.1%—
MATH Level 511.2%—

Knowledge Qwen1.5 4b Chat leads

Gemini 1.0 Pro: 15.6 (#291), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkGemini 1.0 ProQwen1.5 4b Chat
LMArena Expert1059980
GPQA Diamond34%—
MMLU70%—

Multilingual Gemini 1.0 Pro leads

Gemini 1.0 Pro: 33.4 (#252), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkGemini 1.0 ProQwen1.5 4b Chat
LMArena Non-English1138979
LMArena Chinese11241024
LMArena German1125902
LMArena Russian1186952
LMArena French1145—
LMArena Japanese1023—
LMArena Spanish1119—

Instruction Following Gemini 1.0 Pro leads

Gemini 1.0 Pro: 57.6 (#267), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkGemini 1.0 ProQwen1.5 4b Chat
LMArena Instruction Following1114978

Long Context Gemini 1.0 Pro leads

Gemini 1.0 Pro: 34.3 (#249), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkGemini 1.0 ProQwen1.5 4b Chat
LMArena Longer Query1132988

Writing & Preference Gemini 1.0 Pro leads

Gemini 1.0 Pro: 36.0 (#264), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkGemini 1.0 ProQwen1.5 4b Chat
LMArena Text1149997
LMArena Creative Writing1131969
LMArena Multi-Turn1139977

Frequently asked questions

Is Gemini 1.0 Pro better than Qwen1.5 4b Chat?

Qwen1.5 4b Chat is the stronger model overall, scoring 28.8 to 27.3 on the Noometry Index.

Is Gemini 1.0 Pro or Qwen1.5 4b Chat better for coding?

Gemini 1.0 Pro scores higher on coding benchmarks: 32.2 versus 29.1 in the Noometry coding category.

How many benchmarks do Gemini 1.0 Pro and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Gemini 1.0 Pro has 24 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper