Model comparison

Gemini 2.5 Flash-Lite vs Qwen3.7 Plus

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 37.0 on the Noometry Index. Gemini 2.5 Flash-Lite costs 4.0× less per token, which makes it the better buy when Qwen3.7 Plus's lead doesn't matter for your workload.

Last verified . 21 shared benchmarks.

Gemini 2.5 Flash-Lite Google

37.0

Rank #211 Confirmed

Qwen3.7 Plus Alibaba (Qwen)

45.3

Rank #72 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Gemini 2.5 Flash-Lite scores higher in 2 categories and Qwen3.7 Plus in 8 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.7 Plus leads 54.9 to 32.5.
  • The biggest single-benchmark swing is DTBench: 62.8% for Gemini 2.5 Flash-Lite and 84% for Qwen3.7 Plus.
  • Gemini 2.5 Flash-Lite is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.40 / $1.60 for Qwen3.7 Plus.
  • Gemini 2.5 Flash-Lite accepts more context: 1.05M tokens versus 1M.

Side by side

Gemini 2.5 Flash-Lite and Qwen3.7 Plus specifications
Gemini 2.5 Flash-LiteQwen3.7 Plus
ProviderGoogleAlibaba (Qwen)
Noometry Index37.045.3
Released2025-06-172026-06-02
WeightsProprietaryProprietary
Context window1.05M1M
Max output66K131K
Input $ / M tokens$0.10$0.40
Output $ / M tokens$0.40$1.60
Results tracked3332

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 38.5 (#173), Qwen3.7 Plus: 36.6 (#206)

Coding benchmarks
BenchmarkGemini 2.5 Flash-LiteQwen3.7 Plus
LMArena Coding13731473
FrontierCode—10.2%
SciCode—45.5%
WeirdML35.2%—
ALE-Bench325.9—

Agentic & Tool Use Gemini 2.5 Flash-Lite leads

Gemini 2.5 Flash-Lite: 28.0 (#96), Qwen3.7 Plus: 21.4 (#138)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 Flash-LiteQwen3.7 Plus
Berkeley Function Calling Leaderboard36.9%—
OSWorld 2.0—2.8%

Reasoning Qwen3.7 Plus leads

Gemini 2.5 Flash-Lite: 22.2 (#205), Qwen3.7 Plus: 39.3 (#59)

Reasoning benchmarks
BenchmarkGemini 2.5 Flash-LiteQwen3.7 Plus
LMArena Hard Prompts13771460
DTBench62.8%84%
LMCA18.1%37.6%
Epoch Capabilities Index133.94147.37
Kagi LLM Benchmark40.5%—
NYT Connections (extended)—74.8%
CritPt—9.1%
Chess Puzzles—24%
Mystery Game Puzzles—17%

Math Qwen3.7 Plus leads

Gemini 2.5 Flash-Lite: 38.0 (#144), Qwen3.7 Plus: 50.5 (#56)

Math benchmarks
BenchmarkGemini 2.5 Flash-LiteQwen3.7 Plus
LMArena Math13731466
FrontierMath (Tiers 1-3)—34.4%
OTIS Mock AIME 2024-2025—93.3%
Omni-MATH48%—

Knowledge Qwen3.7 Plus leads

Gemini 2.5 Flash-Lite: 32.5 (#210), Qwen3.7 Plus: 54.9 (#51)

Knowledge benchmarks
BenchmarkGemini 2.5 Flash-LiteQwen3.7 Plus
LMArena Expert13731467
GPQA Diamond—87.9%
MMLU-Pro53.7%—
Vectara Hallucination Rate3.3%—
GPQA (HELM)30.9%—

Multimodal Qwen3.7 Plus leads

Gemini 2.5 Flash-Lite: 29.1 (#114), Qwen3.7 Plus: 41.8 (#33)

Multimodal benchmarks
BenchmarkGemini 2.5 Flash-LiteQwen3.7 Plus
LMArena Vision11981279
VPCT30%—
LMArena Document—1444

Multilingual Qwen3.7 Plus leads

Gemini 2.5 Flash-Lite: 49.3 (#134), Qwen3.7 Plus: 54.8 (#38)

Multilingual benchmarks
BenchmarkGemini 2.5 Flash-LiteQwen3.7 Plus
LMArena Non-English13691445
LMArena Chinese14041510
LMArena French13881473
LMArena German13891471
LMArena Japanese13591413
LMArena Korean13601415
LMArena Russian13731457
LMArena Spanish13961457

Instruction Following Qwen3.7 Plus leads

Gemini 2.5 Flash-Lite: 70.0 (#168), Qwen3.7 Plus: 75.8 (#52)

Instruction Following benchmarks
BenchmarkGemini 2.5 Flash-LiteQwen3.7 Plus
LMArena Instruction Following13671440
IFEval81%—

Long Context Qwen3.7 Plus leads

Gemini 2.5 Flash-Lite: 33.3 (#262), Qwen3.7 Plus: 44.5 (#65)

Long Context benchmarks
BenchmarkGemini 2.5 Flash-LiteQwen3.7 Plus
LMArena Longer Query13731455
Fiction.LiveBench47.2%—

Writing & Preference Qwen3.7 Plus leads

Gemini 2.5 Flash-Lite: 56.8 (#135), Qwen3.7 Plus: 64.3 (#56)

Writing & Preference benchmarks
BenchmarkGemini 2.5 Flash-LiteQwen3.7 Plus
LMArena Text13791455
LMArena Creative Writing13671439
LMArena Multi-Turn13661460
WildBench81.8%—

Frequently asked questions

Is Gemini 2.5 Flash-Lite better than Qwen3.7 Plus?

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 37.0 on the Noometry Index. Gemini 2.5 Flash-Lite costs 4.0× less per token, which makes it the better buy when Qwen3.7 Plus's lead doesn't matter for your workload.

Which is cheaper, Gemini 2.5 Flash-Lite or Qwen3.7 Plus?

Gemini 2.5 Flash-Lite is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; Qwen3.7 Plus lists at $0.40 and $1.60.

Is Gemini 2.5 Flash-Lite or Qwen3.7 Plus better for coding?

Gemini 2.5 Flash-Lite scores higher on coding benchmarks: 38.5 versus 36.6 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash-Lite does, with 1.05M tokens against 1M.

How many benchmarks do Gemini 2.5 Flash-Lite and Qwen3.7 Plus share?

21 benchmarks have published results for both models. Gemini 2.5 Flash-Lite has 33 scored results on Noometry and Qwen3.7 Plus has 32.

Related comparisons

Go deeper