Model comparison

Qwen Max vs Yi-Lightning

Yi-Lightning is the stronger model overall, scoring 37.1 to 34.7 on the Noometry Index.

Last verified . 18 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Yi-Lightning 01.AI

37.1

Rank #209 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Qwen Max scores higher in 2 categories and Yi-Lightning in 6 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Yi-Lightning leads 36.2 to 22.3.
  • The biggest single-benchmark swing is Aider Polyglot: 21.8% for Qwen Max and 12.9% for Yi-Lightning.

Side by side

Qwen Max and Yi-Lightning specifications
Qwen MaxYi-Lightning
ProviderAlibaba (Qwen)01.AI
Noometry Index34.737.1
Released2024-04-032024-12-02
WeightsProprietaryProprietary
Context window33K—
Max output8K—
Input $ / M tokens$1.60—
Output $ / M tokens$6.40—
Results tracked2318

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Max leads

Qwen Max: 30.7 (#292), Yi-Lightning: 27.4 (#320)

Coding benchmarks
BenchmarkQwen MaxYi-Lightning
Aider Polyglot21.8%12.9%
LMArena Coding12881312

Reasoning Too close to call

Qwen Max: 25.1 (#151), Yi-Lightning: 25.9 (#139)

Reasoning benchmarks
BenchmarkQwen MaxYi-Lightning
LMArena Hard Prompts12691302

Math Yi-Lightning leads

Qwen Max: 22.3 (#276), Yi-Lightning: 36.2 (#172)

Math benchmarks
BenchmarkQwen MaxYi-Lightning
LMArena Math12751300
OTIS Mock AIME 2024-202516.1%—
MATH Level 567.2%—
FrontierMath (Feb 2025 set)1%—

Knowledge Yi-Lightning leads

Qwen Max: 30.3 (#228), Yi-Lightning: 35.4 (#185)

Knowledge benchmarks
BenchmarkQwen MaxYi-Lightning
LMArena Expert12481286
GPQA Diamond56.1%—

Multilingual Too close to call

Qwen Max: 41.8 (#202), Yi-Lightning: 42.2 (#196)

Multilingual benchmarks
BenchmarkQwen MaxYi-Lightning
LMArena Non-English12631269
LMArena Chinese12541320
LMArena French13301305
LMArena German12541267
LMArena Japanese12051228
LMArena Korean11421192
LMArena Russian12741255
LMArena Spanish12901315

Instruction Following Too close to call

Qwen Max: 66.5 (#208), Yi-Lightning: 67.4 (#196)

Instruction Following benchmarks
BenchmarkQwen MaxYi-Lightning
LMArena Instruction Following12621278

Long Context Too close to call

Qwen Max: 39.4 (#180), Yi-Lightning: 39.4 (#181)

Long Context benchmarks
BenchmarkQwen MaxYi-Lightning
LMArena Longer Query12881297
Fiction.LiveBench66.7%—

Writing & Preference Yi-Lightning leads

Qwen Max: 47.8 (#205), Yi-Lightning: 50.2 (#184)

Writing & Preference benchmarks
BenchmarkQwen MaxYi-Lightning
LMArena Text12821302
LMArena Creative Writing12481280
LMArena Multi-Turn12771311

Frequently asked questions

Is Qwen Max better than Yi-Lightning?

Yi-Lightning is the stronger model overall, scoring 37.1 to 34.7 on the Noometry Index.

Is Qwen Max or Yi-Lightning better for coding?

Qwen Max scores higher on coding benchmarks: 30.7 versus 27.4 in the Noometry coding category.

How many benchmarks do Qwen Max and Yi-Lightning share?

18 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Yi-Lightning has 18.

Related comparisons

Go deeper