Model comparison

Nova 2 Lite vs Qwen3.7 Plus

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 39.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Nova 2 Lite Amazon

39.7

Rank #161 Confirmed

Qwen3.7 Plus Alibaba (Qwen)

45.3

Rank #72 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Nova 2 Lite scores higher in 2 categories and Qwen3.7 Plus in 7 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.7 Plus leads 50.5 to 37.5.
  • Qwen3.7 Plus is cheaper at $0.40 / $1.60 per million input/output tokens, against $0.30 / $2.50 for Nova 2 Lite.

Side by side

Nova 2 Lite and Qwen3.7 Plus specifications
Nova 2 LiteQwen3.7 Plus
ProviderAmazonAlibaba (Qwen)
Noometry Index39.745.3
Released2025-12-012026-06-02
WeightsProprietaryProprietary
Context window1M1M
Max output64K131K
Input $ / M tokens$0.30$0.40
Output $ / M tokens$2.50$1.60
Results tracked1932

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Nova 2 Lite leads

Nova 2 Lite: 40.7 (#134), Qwen3.7 Plus: 36.6 (#206)

Coding benchmarks
BenchmarkNova 2 LiteQwen3.7 Plus
LMArena Coding13851473
FrontierCode—10.2%
SciCode—45.5%

Agentic & Tool Use Nova 2 Lite leads

Nova 2 Lite: 24.1 (#121), Qwen3.7 Plus: 21.4 (#138)

Agentic & Tool Use benchmarks
BenchmarkNova 2 LiteQwen3.7 Plus
Berkeley Function Calling Leaderboard27.1%—
OSWorld 2.0—2.8%

Reasoning Qwen3.7 Plus leads

Nova 2 Lite: 27.5 (#118), Qwen3.7 Plus: 39.3 (#59)

Reasoning benchmarks
BenchmarkNova 2 LiteQwen3.7 Plus
LMArena Hard Prompts13641460
NYT Connections (extended)—74.8%
CritPt—9.1%
Chess Puzzles—24%
Mystery Game Puzzles—17%
DTBench—84%
LMCA—37.6%
Epoch Capabilities Index—147.37

Math Qwen3.7 Plus leads

Nova 2 Lite: 37.5 (#156), Qwen3.7 Plus: 50.5 (#56)

Math benchmarks
BenchmarkNova 2 LiteQwen3.7 Plus
LMArena Math13591466
FrontierMath (Tiers 1-3)—34.4%
OTIS Mock AIME 2024-2025—93.3%

Knowledge Qwen3.7 Plus leads

Nova 2 Lite: 43.0 (#94), Qwen3.7 Plus: 54.9 (#51)

Knowledge benchmarks
BenchmarkNova 2 LiteQwen3.7 Plus
LMArena Expert13581467
GPQA Diamond—87.9%
Vectara Hallucination Rate5.1%—

Multimodal Not comparable

Nova 2 Lite: —, Qwen3.7 Plus: 41.8 (#33)

Multimodal benchmarks
BenchmarkNova 2 LiteQwen3.7 Plus
LMArena Vision—1279
LMArena Document—1444

Multilingual Qwen3.7 Plus leads

Nova 2 Lite: 47.1 (#153), Qwen3.7 Plus: 54.8 (#38)

Multilingual benchmarks
BenchmarkNova 2 LiteQwen3.7 Plus
LMArena Non-English13371445
LMArena Chinese13641510
LMArena French13811473
LMArena German13431471
LMArena Japanese12711413
LMArena Korean12841415
LMArena Russian13431457
LMArena Spanish13731457

Instruction Following Qwen3.7 Plus leads

Nova 2 Lite: 70.5 (#161), Qwen3.7 Plus: 75.8 (#52)

Instruction Following benchmarks
BenchmarkNova 2 LiteQwen3.7 Plus
LMArena Instruction Following13351440

Long Context Qwen3.7 Plus leads

Nova 2 Lite: 40.6 (#150), Qwen3.7 Plus: 44.5 (#65)

Long Context benchmarks
BenchmarkNova 2 LiteQwen3.7 Plus
LMArena Longer Query13351455

Writing & Preference Qwen3.7 Plus leads

Nova 2 Lite: 53.9 (#154), Qwen3.7 Plus: 64.3 (#56)

Writing & Preference benchmarks
BenchmarkNova 2 LiteQwen3.7 Plus
LMArena Text13621455
LMArena Creative Writing12911439
LMArena Multi-Turn13381460

Frequently asked questions

Is Nova 2 Lite better than Qwen3.7 Plus?

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 39.7 on the Noometry Index.

Which is cheaper, Nova 2 Lite or Qwen3.7 Plus?

Qwen3.7 Plus is cheaper. It lists at $0.40 per million input tokens and $1.60 per million output tokens; Nova 2 Lite lists at $0.30 and $2.50.

Is Nova 2 Lite or Qwen3.7 Plus better for coding?

Nova 2 Lite scores higher on coding benchmarks: 40.7 versus 36.6 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do Nova 2 Lite and Qwen3.7 Plus share?

17 benchmarks have published results for both models. Nova 2 Lite has 19 scored results on Noometry and Qwen3.7 Plus has 32.

Related comparisons

Go deeper