Model comparison

Llama 4 Scout vs Yi-Lightning

Yi-Lightning is the stronger model overall, scoring 37.1 to 27.7 on the Noometry Index.

Last verified . 17 shared benchmarks.

Llama 4 Scout Meta

27.7

Rank #330 Confirmed

Yi-Lightning 01.AI

37.1

Rank #209 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Llama 4 Scout scores higher in 0 categories and Yi-Lightning in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Yi-Lightning leads 25.9 to 9.1.
  • Llama 4 Scout has downloadable open weights; the other is API-only.

Side by side

Llama 4 Scout and Yi-Lightning specifications
Llama 4 ScoutYi-Lightning
ProviderMeta01.AI
Noometry Index27.737.1
Released2025-04-052024-12-02
WeightsOpenProprietary
Context window128K—
Max output4K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.30—
Results tracked4318

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Yi-Lightning leads

Llama 4 Scout: 20.2 (#339), Yi-Lightning: 27.4 (#320)

Coding benchmarks
BenchmarkLlama 4 ScoutYi-Lightning
LMArena Coding12861312
SWE-bench Verified (bash only)9.1%—
Aider Polyglot—12.9%
SciCode17%—
BigCodeBench Complete43.1%—

Agentic & Tool Use Not comparable

Llama 4 Scout: 24.6 (#119), Yi-Lightning: —

Agentic & Tool Use benchmarks
BenchmarkLlama 4 ScoutYi-Lightning
Berkeley Function Calling Leaderboard28.1%—

Reasoning Yi-Lightning leads

Llama 4 Scout: 9.1 (#345), Yi-Lightning: 25.9 (#139)

Reasoning benchmarks
BenchmarkLlama 4 ScoutYi-Lightning
LMArena Hard Prompts12661302
ARC-AGI-20%—
Kagi LLM Benchmark36.9%—
ARC-AGI-10.5%—
CritPt0%—
DTBench57.9%—
LMCA12%—
Epoch Capabilities Index129.64—
ForecastBench57.5—

Math Yi-Lightning leads

Llama 4 Scout: 19.6 (#286), Yi-Lightning: 36.2 (#172)

Math benchmarks
BenchmarkLlama 4 ScoutYi-Lightning
LMArena Math12871300
OTIS Mock AIME 2024-20257.8%—
Omni-MATH37.3%—
MATH Level 562.3%—
FrontierMath (Feb 2025 set)0%—

Knowledge Yi-Lightning leads

Llama 4 Scout: 31.9 (#217), Yi-Lightning: 35.4 (#185)

Knowledge benchmarks
BenchmarkLlama 4 ScoutYi-Lightning
LMArena Expert12351286
GPQA Diamond51.8%—
MMLU-Pro74.2%—
Vectara Hallucination Rate7.7%—
GPQA (HELM)50.7%—

Multimodal Not comparable

Llama 4 Scout: 32.2 (#102), Yi-Lightning: —

Multimodal benchmarks
BenchmarkLlama 4 ScoutYi-Lightning
LMArena Vision1118—
SpatialViz-Bench34.2%—

Multilingual Yi-Lightning leads

Llama 4 Scout: 41.0 (#212), Yi-Lightning: 42.2 (#196)

Multilingual benchmarks
BenchmarkLlama 4 ScoutYi-Lightning
LMArena Non-English12521269
LMArena Chinese12551320
LMArena French12821305
LMArena German12721267
LMArena Japanese12061228
LMArena Korean12071192
LMArena Russian12631255
LMArena Spanish12781315

Instruction Following Yi-Lightning leads

Llama 4 Scout: 65.8 (#217), Yi-Lightning: 67.4 (#196)

Instruction Following benchmarks
BenchmarkLlama 4 ScoutYi-Lightning
LMArena Instruction Following12481278
IFEval81.8%—

Long Context Yi-Lightning leads

Llama 4 Scout: 27.5 (#294), Yi-Lightning: 39.4 (#181)

Long Context benchmarks
BenchmarkLlama 4 ScoutYi-Lightning
LMArena Longer Query12651297
Fiction.LiveBench36%—

Writing & Preference Yi-Lightning leads

Llama 4 Scout: 37.0 (#261), Yi-Lightning: 50.2 (#184)

Writing & Preference benchmarks
BenchmarkLlama 4 ScoutYi-Lightning
LMArena Text12791302
LMArena Creative Writing12491280
LMArena Multi-Turn12801311
EQ-Bench Creative Writing783—
WildBench78%—

Frequently asked questions

Is Llama 4 Scout better than Yi-Lightning?

Yi-Lightning is the stronger model overall, scoring 37.1 to 27.7 on the Noometry Index.

Is Llama 4 Scout or Yi-Lightning better for coding?

Yi-Lightning scores higher on coding benchmarks: 27.4 versus 20.2 in the Noometry coding category.

How many benchmarks do Llama 4 Scout and Yi-Lightning share?

17 benchmarks have published results for both models. Llama 4 Scout has 43 scored results on Noometry and Yi-Lightning has 18.

Related comparisons

Go deeper