Model comparison

Gemini 2.0 Flash (Feb 2025) vs Qwen Max

Gemini 2.0 Flash (Feb 2025) and Qwen Max score almost the same on the Noometry Index (35.1 vs 34.7), so choose on price, context window or the category you care about most.

Last verified . 23 shared benchmarks.

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Gemini 2.0 Flash (Feb 2025) scores higher in 5 categories and Qwen Max in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemini 2.0 Flash (Feb 2025) leads 37.9 to 22.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 57.8% for Gemini 2.0 Flash (Feb 2025) and 16.1% for Qwen Max.

Side by side

Gemini 2.0 Flash (Feb 2025) and Qwen Max specifications
Gemini 2.0 Flash (Feb 2025)Qwen Max
ProviderGoogleAlibaba (Qwen)
Noometry Index35.134.7
Released2024-12-062024-04-03
WeightsProprietaryProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked5423

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Max leads

Gemini 2.0 Flash (Feb 2025): 28.4 (#315), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Max
Aider Polyglot38.2%21.8%
LMArena Coding13501288
SWE-bench Verified (bash only)13.5%—
WeirdML25.8%—
BigCodeBench Instruct45.9%—
LiveBench Coding63.4%—
BigCodeBench Complete59.9%—
CadEval30%—

Agentic & Tool Use Not comparable

Gemini 2.0 Flash (Feb 2025): 28.1 (#92), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Max
TheAgentCompany11.4%—

Reasoning Qwen Max leads

Gemini 2.0 Flash (Feb 2025): 15.2 (#318), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Max
LMArena Hard Prompts13461269
ARC-AGI-21.3%—
SimpleBench31.1%—
Kagi LLM Benchmark37.8%—
EnigmaEval1.1%—
LiveBench Reasoning78.2%—
DTBench63.2%—
LiveBench Data Analysis69.4%—
Epoch Capabilities Index135.36—
LiveBench66.9%—

Math Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 37.9 (#146), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Max
OTIS Mock AIME 2024-202557.8%16.1%
LMArena Math13521275
MATH Level 582.2%67.2%
FrontierMath (Feb 2025 set)1.7%1%
Omni-MATH45.9%—
LiveBench Math75.8%—

Knowledge Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 32.0 (#213), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Max
GPQA Diamond64.1%56.1%
LMArena Expert13391248
Humanity's Last Exam6.6%—
MMLU-Pro73.7%—
Confabulations12.4%—
GPQA (HELM)55.6%—
MMLU79.7%—

Multimodal Not comparable

Gemini 2.0 Flash (Feb 2025): 36.5 (#79), Qwen Max: —

Multimodal benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Max
LMArena Vision1158—
GeoBench77%—

Multilingual Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 47.4 (#149), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Max
LMArena Non-English13421263
LMArena Chinese13731254
LMArena French13911330
LMArena German13531254
LMArena Japanese12941205
LMArena Korean13131142
LMArena Russian13511274
LMArena Spanish13631290

Instruction Following Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 74.4 (#97), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Max
LMArena Instruction Following13361262
LiveBench Instruction Following85.8%—
IFEval84.1%—

Long Context Qwen Max leads

Gemini 2.0 Flash (Feb 2025): 38.1 (#203), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Max
Fiction.LiveBench61.1%66.7%
LMArena Longer Query13441288

Writing & Preference Gemini 2.0 Flash (Feb 2025) leads

Gemini 2.0 Flash (Feb 2025): 49.5 (#190), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkGemini 2.0 Flash (Feb 2025)Qwen Max
LMArena Text13541282
LMArena Creative Writing13401248
LMArena Multi-Turn13501277
Short-Story Creative Writing73.8%—
EQ-Bench Creative Writing1128—
WildBench80%—
LiveBench Language51.3%—

Frequently asked questions

Is Gemini 2.0 Flash (Feb 2025) better than Qwen Max?

Gemini 2.0 Flash (Feb 2025) and Qwen Max score almost the same on the Noometry Index (35.1 vs 34.7), so choose on price, context window or the category you care about most.

Is Gemini 2.0 Flash (Feb 2025) or Qwen Max better for coding?

Qwen Max scores higher on coding benchmarks: 30.7 versus 28.4 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Flash (Feb 2025) and Qwen Max share?

23 benchmarks have published results for both models. Gemini 2.0 Flash (Feb 2025) has 54 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper