Model comparison

Olmo 3.1 32b Think vs Yi-Lightning

Olmo 3.1 32b Think and Yi-Lightning score almost the same on the Noometry Index (37.9 vs 37.1), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Yi-Lightning 01.AI

37.1

Rank #209 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Olmo 3.1 32b Think scores higher in 3 categories and Yi-Lightning in 5 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Olmo 3.1 32b Think leads 37.7 to 27.4.
  • Olmo 3.1 32b Think has downloadable open weights; the other is API-only.

Side by side

Olmo 3.1 32b Think and Yi-Lightning specifications
Olmo 3.1 32b ThinkYi-Lightning
ProviderAllen Institute for AI (Ai2)01.AI
Noometry Index37.937.1
Released—2024-12-02
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1518

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 37.7 (#189), Yi-Lightning: 27.4 (#320)

Coding benchmarks
BenchmarkOlmo 3.1 32b ThinkYi-Lightning
LMArena Coding12911312
Aider Polyglot—12.9%

Reasoning Too close to call

Olmo 3.1 32b Think: 25.2 (#150), Yi-Lightning: 25.9 (#139)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b ThinkYi-Lightning
LMArena Hard Prompts12721302

Math Too close to call

Olmo 3.1 32b Think: 36.3 (#168), Yi-Lightning: 36.2 (#172)

Math benchmarks
BenchmarkOlmo 3.1 32b ThinkYi-Lightning
LMArena Math13051300

Knowledge Too close to call

Olmo 3.1 32b Think: 35.7 (#181), Yi-Lightning: 35.4 (#185)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b ThinkYi-Lightning
LMArena Expert12951286

Multilingual Yi-Lightning leads

Olmo 3.1 32b Think: 38.1 (#231), Yi-Lightning: 42.2 (#196)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b ThinkYi-Lightning
LMArena Non-English12091269
LMArena Chinese12421320
LMArena French12601305
LMArena German12621267
LMArena Russian11931255
LMArena Spanish12891315
LMArena Japanese—1228
LMArena Korean—1192

Instruction Following Yi-Lightning leads

Olmo 3.1 32b Think: 65.6 (#218), Yi-Lightning: 67.4 (#196)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b ThinkYi-Lightning
LMArena Instruction Following12471278

Long Context Too close to call

Olmo 3.1 32b Think: 38.6 (#195), Yi-Lightning: 39.4 (#181)

Long Context benchmarks
BenchmarkOlmo 3.1 32b ThinkYi-Lightning
LMArena Longer Query12721297

Writing & Preference Yi-Lightning leads

Olmo 3.1 32b Think: 46.2 (#220), Yi-Lightning: 50.2 (#184)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b ThinkYi-Lightning
LMArena Text12721302
LMArena Creative Writing12261280
LMArena Multi-Turn12521311

Frequently asked questions

Is Olmo 3.1 32b Think better than Yi-Lightning?

Olmo 3.1 32b Think and Yi-Lightning score almost the same on the Noometry Index (37.9 vs 37.1), so choose on price, context window or the category you care about most.

Is Olmo 3.1 32b Think or Yi-Lightning better for coding?

Olmo 3.1 32b Think scores higher on coding benchmarks: 37.7 versus 27.4 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Think and Yi-Lightning share?

15 benchmarks have published results for both models. Olmo 3.1 32b Think has 15 scored results on Noometry and Yi-Lightning has 18.

Related comparisons

Go deeper