Model comparison

Gemini 2.0 Flash (Feb 2025) vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 35.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemini 2.0 Flash (Feb 2025) scores higher in 0 categories and Qwen3.5 Max Preview in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 Max Preview leads 66.0 to 49.5.

Side by side

Gemini 2.0 Flash (Feb 2025) and Qwen3.5 Max Preview specifications
Gemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
ProviderGoogleAlibaba (Qwen)
Noometry Index35.145.3
Released2024-12-06—
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked5417

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Gemini 2.0 Flash (Feb 2025): 28.4 (#315), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
LMArena Coding13501487
SWE-bench Verified (bash only)13.5%—
Aider Polyglot38.2%—
WeirdML25.8%—
BigCodeBench Instruct45.9%—
LiveBench Coding63.4%—
BigCodeBench Complete59.9%—
CadEval30%—

Agentic & Tool Use Not comparable

Gemini 2.0 Flash (Feb 2025): 28.1 (#92), Qwen3.5 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
TheAgentCompany11.4%—

Reasoning Qwen3.5 Max Preview leads

Gemini 2.0 Flash (Feb 2025): 15.2 (#318), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
LMArena Hard Prompts13461483
ARC-AGI-21.3%—
SimpleBench31.1%—
Kagi LLM Benchmark37.8%—
EnigmaEval1.1%—
LiveBench Reasoning78.2%—
DTBench63.2%—
LiveBench Data Analysis69.4%—
Epoch Capabilities Index135.36—
LiveBench66.9%—

Math Qwen3.5 Max Preview leads

Gemini 2.0 Flash (Feb 2025): 37.9 (#146), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
LMArena Math13521474
OTIS Mock AIME 2024-202557.8%—
Omni-MATH45.9%—
LiveBench Math75.8%—
MATH Level 582.2%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Qwen3.5 Max Preview leads

Gemini 2.0 Flash (Feb 2025): 32.0 (#213), Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
LMArena Expert13391489
GPQA Diamond64.1%—
Humanity's Last Exam6.6%—
MMLU-Pro73.7%—
Confabulations12.4%—
GPQA (HELM)55.6%—
MMLU79.7%—

Multimodal Not comparable

Gemini 2.0 Flash (Feb 2025): 36.5 (#79), Qwen3.5 Max Preview: —

Multimodal benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
LMArena Vision1158—
GeoBench77%—

Multilingual Qwen3.5 Max Preview leads

Gemini 2.0 Flash (Feb 2025): 47.4 (#149), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
LMArena Non-English13421465
LMArena Chinese13731534
LMArena French13911484
LMArena German13531487
LMArena Japanese12941495
LMArena Korean13131438
LMArena Russian13511471
LMArena Spanish13631470

Instruction Following Qwen3.5 Max Preview leads

Gemini 2.0 Flash (Feb 2025): 74.4 (#97), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
LMArena Instruction Following13361467
LiveBench Instruction Following85.8%—
IFEval84.1%—

Long Context Qwen3.5 Max Preview leads

Gemini 2.0 Flash (Feb 2025): 38.1 (#203), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
LMArena Longer Query13441476
Fiction.LiveBench61.1%—

Writing & Preference Qwen3.5 Max Preview leads

Gemini 2.0 Flash (Feb 2025): 49.5 (#190), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen3.5 Max Preview
LMArena Text13541470
LMArena Creative Writing13401464
LMArena Multi-Turn13501478
Short-Story Creative Writing73.8%—
EQ-Bench Creative Writing1128—
WildBench80%—
LiveBench Language51.3%—

Frequently asked questions

Is Gemini 2.0 Flash (Feb 2025) better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 35.1 on the Noometry Index.

Is Gemini 2.0 Flash (Feb 2025) or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 28.4 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Flash (Feb 2025) and Qwen3.5 Max Preview share?

17 benchmarks have published results for both models. Gemini 2.0 Flash (Feb 2025) has 54 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper