Model comparison

Gemini 2.5 Flash vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 39.3 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemini 2.5 Flash Google

39.3

Rank #170 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemini 2.5 Flash scores higher in 1 category and Qwen3.5 Max Preview in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.5 Max Preview leads 30.8 to 18.1.

Side by side

Gemini 2.5 Flash and Qwen3.5 Max Preview specifications
Gemini 2.5 FlashQwen3.5 Max Preview
ProviderGoogleAlibaba (Qwen)
Noometry Index39.345.3
Released2025-04-17—
WeightsProprietaryProprietary
Context window1.05M—
Max output66K—
Input $ / M tokens$0.30—
Output $ / M tokens$2.50—
Results tracked5417

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Gemini 2.5 Flash: 35.8 (#220), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkGemini 2.5 FlashQwen3.5 Max Preview
LMArena Coding14241487
SWE-bench Verified (bash only)28.7%—
Aider Polyglot55.1%—
WeirdML41.9%—
ALE-Bench661.88—

Agentic & Tool Use Not comparable

Gemini 2.5 Flash: 30.8 (#74), Qwen3.5 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 FlashQwen3.5 Max Preview
Terminal-Bench17.1%—
Berkeley Function Calling Leaderboard56.2%—
TheAgentCompany41.1%—
BALROG33.5%—
Vending-Bench 2548.84—

Reasoning Qwen3.5 Max Preview leads

Gemini 2.5 Flash: 18.1 (#286), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkGemini 2.5 FlashQwen3.5 Max Preview
LMArena Hard Prompts14221483
ARC-AGI-22.5%—
SimpleBench41.2%—
Kagi LLM Benchmark56.8%—
ARC-AGI-133.3%—
CritPt1.1%—
EnigmaEval2.7%—
DTBench76.5%—
LMCA27.5%—
Epoch Capabilities Index143.03—
ForecastBench60.6—

Math Too close to call

Gemini 2.5 Flash: 39.9 (#98), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkGemini 2.5 FlashQwen3.5 Max Preview
LMArena Math14151474
OTIS Mock AIME 2024-202573.1%—
Omni-MATH38.5%—
FrontierMath (Feb 2025 set)4.8%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Qwen3.5 Max Preview leads

Gemini 2.5 Flash: 36.4 (#168), Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkGemini 2.5 FlashQwen3.5 Max Preview
LMArena Expert14261489
Humanity's Last Exam12.1%—
MMLU-Pro63.9%—
Confabulations16.8%—
Vectara Hallucination Rate7.8%—
GPQA (HELM)39%—

Multimodal Not comparable

Gemini 2.5 Flash: 41.8 (#32), Qwen3.5 Max Preview: —

Multimodal benchmarks
BenchmarkGemini 2.5 FlashQwen3.5 Max Preview
LMArena Vision1253—
GeoBench76%—
VPCT46.2%—
SpatialViz-Bench36.9%—

Multilingual Qwen3.5 Max Preview leads

Gemini 2.5 Flash: 52.3 (#88), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkGemini 2.5 FlashQwen3.5 Max Preview
LMArena Non-English14091465
LMArena Chinese14501534
LMArena French14331484
LMArena German14181487
LMArena Japanese14051495
LMArena Korean13851438
LMArena Russian14151471
LMArena Spanish14211470

Instruction Following Qwen3.5 Max Preview leads

Gemini 2.5 Flash: 75.7 (#54), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkGemini 2.5 FlashQwen3.5 Max Preview
LMArena Instruction Following14051467
IFEval89.8%—

Long Context Gemini 2.5 Flash leads

Gemini 2.5 Flash: 47.5 (#17), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkGemini 2.5 FlashQwen3.5 Max Preview
LMArena Longer Query14191476
Fiction.LiveBench77.8%—

Writing & Preference Qwen3.5 Max Preview leads

Gemini 2.5 Flash: 53.8 (#157), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkGemini 2.5 FlashQwen3.5 Max Preview
LMArena Text14171470
LMArena Creative Writing14001464
LMArena Multi-Turn14081478
Short-Story Creative Writing76.5%—
EQ-Bench Creative Writing1137—
WildBench81.7%—

Frequently asked questions

Is Gemini 2.5 Flash better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 39.3 on the Noometry Index.

Is Gemini 2.5 Flash or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 35.8 in the Noometry coding category.

How many benchmarks do Gemini 2.5 Flash and Qwen3.5 Max Preview share?

17 benchmarks have published results for both models. Gemini 2.5 Flash has 54 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper