Model comparison

QwQ-32B vs Yi-Lightning

QwQ-32B is the stronger model overall, scoring 39.8 to 37.1 on the Noometry Index.

Last verified . 18 shared benchmarks.

QwQ-32B Alibaba (Qwen)

39.8

Rank #159 Confirmed

Yi-Lightning 01.AI

37.1

Rank #209 Confirmed

Summary

  • They share 18 benchmarks with published results for both. QwQ-32B scores higher in 7 categories and Yi-Lightning in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where QwQ-32B leads 49.0 to 39.4.
  • The biggest single-benchmark swing is Aider Polyglot: 20.9% for QwQ-32B and 12.9% for Yi-Lightning.
  • QwQ-32B has downloadable open weights; the other is API-only.

Side by side

QwQ-32B and Yi-Lightning specifications
QwQ-32BYi-Lightning
ProviderAlibaba (Qwen)01.AI
Noometry Index39.837.1
Released2024-11-282024-12-02
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3618

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding QwQ-32B leads

QwQ-32B: 35.4 (#226), Yi-Lightning: 27.4 (#320)

Coding benchmarks
BenchmarkQwQ-32BYi-Lightning
Aider Polyglot20.9%12.9%
LMArena Coding13331312
BigCodeBench Instruct44.6%—
LiveBench Coding72.2%—
BigCodeBench Complete54.4%—

Reasoning Yi-Lightning leads

QwQ-32B: 23.7 (#174), Yi-Lightning: 25.9 (#139)

Reasoning benchmarks
BenchmarkQwQ-32BYi-Lightning
LMArena Hard Prompts13251302
Chess Puzzles5%—
LiveBench Reasoning83.5%—
LiveBench Data Analysis65%—
Epoch Capabilities Index137.6—
ForecastBench58.3—
LiveBench72%—

Math QwQ-32B leads

QwQ-32B: 38.0 (#143), Yi-Lightning: 36.2 (#172)

Math benchmarks
BenchmarkQwQ-32BYi-Lightning
LMArena Math13591300
OTIS Mock AIME 2024-202559.2%—
LiveBench Math77.8%—

Knowledge QwQ-32B leads

QwQ-32B: 37.2 (#158), Yi-Lightning: 35.4 (#185)

Knowledge benchmarks
BenchmarkQwQ-32BYi-Lightning
LMArena Expert13241286
GPQA Diamond65.3%—
Confabulations15.6%—

Multilingual QwQ-32B leads

QwQ-32B: 44.8 (#176), Yi-Lightning: 42.2 (#196)

Multilingual benchmarks
BenchmarkQwQ-32BYi-Lightning
LMArena Non-English13051269
LMArena Chinese13781320
LMArena French13361305
LMArena German13131267
LMArena Japanese12621228
LMArena Korean12791192
LMArena Russian12971255
LMArena Spanish13541315

Instruction Following QwQ-32B leads

QwQ-32B: 72.6 (#137), Yi-Lightning: 67.4 (#196)

Instruction Following benchmarks
BenchmarkQwQ-32BYi-Lightning
LMArena Instruction Following12971278
LiveBench Instruction Following81.8%—

Long Context QwQ-32B leads

QwQ-32B: 49.0 (#11), Yi-Lightning: 39.4 (#181)

Long Context benchmarks
BenchmarkQwQ-32BYi-Lightning
LMArena Longer Query13081297
Fiction.LiveBench83.3%—

Writing & Preference Too close to call

QwQ-32B: 50.6 (#180), Yi-Lightning: 50.2 (#184)

Writing & Preference benchmarks
BenchmarkQwQ-32BYi-Lightning
LMArena Text13291302
LMArena Creative Writing12881280
LMArena Multi-Turn13141311
Short-Story Creative Writing80.2%—
EQ-Bench Creative Writing1257—
LiveBench Language51.4%—

Frequently asked questions

Is QwQ-32B better than Yi-Lightning?

QwQ-32B is the stronger model overall, scoring 39.8 to 37.1 on the Noometry Index.

Is QwQ-32B or Yi-Lightning better for coding?

QwQ-32B scores higher on coding benchmarks: 35.4 versus 27.4 in the Noometry coding category.

How many benchmarks do QwQ-32B and Yi-Lightning share?

18 benchmarks have published results for both models. QwQ-32B has 36 scored results on Noometry and Yi-Lightning has 18.

Related comparisons

Go deeper