Model comparison

Gemini 1.5 Flash 8B vs Qwen1.5 4b Chat

Gemini 1.5 Flash 8B is the stronger model overall, scoring 29.9 to 28.8 on the Noometry Index.

Last verified . 13 shared benchmarks.

Gemini 1.5 Flash 8B Google

29.9

Rank #301 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Gemini 1.5 Flash 8B scores higher in 6 categories and Qwen1.5 4b Chat in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemini 1.5 Flash 8B leads 42.8 to 23.8.
  • Qwen1.5 4b Chat has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Flash 8B and Qwen1.5 4b Chat specifications
Gemini 1.5 Flash 8BQwen1.5 4b Chat
ProviderGoogleAlibaba (Qwen)
Noometry Index29.928.8
Released2024-10-03—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2113

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 35.5 (#225), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkGemini 1.5 Flash 8BQwen1.5 4b Chat
LMArena Coding1218999

Reasoning Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 20.0 (#244), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash 8BQwen1.5 4b Chat
LMArena Hard Prompts1209976
DTBench50%—

Math Qwen1.5 4b Chat leads

Gemini 1.5 Flash 8B: 14.2 (#302), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkGemini 1.5 Flash 8BQwen1.5 4b Chat
LMArena Math12071026
OTIS Mock AIME 2024-20254.6%—

Knowledge Qwen1.5 4b Chat leads

Gemini 1.5 Flash 8B: 16.0 (#289), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash 8BQwen1.5 4b Chat
LMArena Expert1185980
GPQA Diamond33%—

Multimodal Not comparable

Gemini 1.5 Flash 8B: 28.2 (#115), Qwen1.5 4b Chat: —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash 8BQwen1.5 4b Chat
LMArena Vision1044—

Multilingual Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 38.5 (#229), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash 8BQwen1.5 4b Chat
LMArena Non-English1215979
LMArena Chinese12311024
LMArena German1206902
LMArena Russian1236952
LMArena French1234—
LMArena Japanese1150—
LMArena Korean1140—
LMArena Spanish1212—

Instruction Following Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 62.8 (#236), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash 8BQwen1.5 4b Chat
LMArena Instruction Following1199978

Long Context Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 37.0 (#225), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkGemini 1.5 Flash 8BQwen1.5 4b Chat
LMArena Longer Query1219988

Writing & Preference Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 42.8 (#232), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash 8BQwen1.5 4b Chat
LMArena Text1226997
LMArena Creative Writing1218969
LMArena Multi-Turn1185977

Frequently asked questions

Is Gemini 1.5 Flash 8B better than Qwen1.5 4b Chat?

Gemini 1.5 Flash 8B is the stronger model overall, scoring 29.9 to 28.8 on the Noometry Index.

Is Gemini 1.5 Flash 8B or Qwen1.5 4b Chat better for coding?

Gemini 1.5 Flash 8B scores higher on coding benchmarks: 35.5 versus 29.1 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash 8B and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Gemini 1.5 Flash 8B has 21 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper