Model comparison

GPT-4 vs Yi-Lightning

Yi-Lightning is the stronger model overall, scoring 37.1 to 29.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Yi-Lightning 01.AI

37.1

Rank #209 Confirmed

Summary

  • They share 17 benchmarks with published results for both. GPT-4 scores higher in 1 category and Yi-Lightning in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Yi-Lightning leads 36.2 to 10.8.

Side by side

GPT-4 and Yi-Lightning specifications
GPT-4Yi-Lightning
ProviderOpenAI01.AI
Noometry Index29.137.1
Released2023-03-142024-12-02
WeightsProprietaryProprietary
Context window8K—
Max output8K—
Input $ / M tokens$30—
Output $ / M tokens$60—
Results tracked3818

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4 leads

GPT-4: 31.6 (#283), Yi-Lightning: 27.4 (#320)

Coding benchmarks
BenchmarkGPT-4Yi-Lightning
LMArena Coding12541312
Aider Polyglot—12.9%
WeirdML12.4%—
BigCodeBench Instruct46%—
BigCodeBench Complete57.2%—
HumanEval+79.3%—

Agentic & Tool Use Not comparable

GPT-4: —, Yi-Lightning: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4Yi-Lightning
METR Time Horizons36.1%—

Reasoning Yi-Lightning leads

GPT-4: 17.8 (#289), Yi-Lightning: 25.9 (#139)

Reasoning benchmarks
BenchmarkGPT-4Yi-Lightning
LMArena Hard Prompts12411302
Chess Puzzles4%—
Mystery Game Puzzles12%—
DTBench62.7%—
LMCA17.1%—
BIG-Bench Hard75.1%—
Epoch Capabilities Index125.89—
ForecastBench57.8—
HellaSwag95.3%—
WinoGrande87.5%—

Math Yi-Lightning leads

GPT-4: 10.8 (#309), Yi-Lightning: 36.2 (#172)

Math benchmarks
BenchmarkGPT-4Yi-Lightning
LMArena Math12691300
OTIS Mock AIME 2024-20251.1%—
MATH Level 523%—
GSM8K92%—

Knowledge Yi-Lightning leads

GPT-4: 18.4 (#282), Yi-Lightning: 35.4 (#185)

Knowledge benchmarks
BenchmarkGPT-4Yi-Lightning
LMArena Expert12111286
GPQA Diamond35.7%—
MMLU86.4%—
TriviaQA84.8%—

Multilingual Yi-Lightning leads

GPT-4: 40.6 (#215), Yi-Lightning: 42.2 (#196)

Multilingual benchmarks
BenchmarkGPT-4Yi-Lightning
LMArena Non-English12461269
LMArena Chinese12421320
LMArena French12831305
LMArena German12511267
LMArena Japanese12091228
LMArena Korean11841192
LMArena Russian12511255
LMArena Spanish12611315

Instruction Following Yi-Lightning leads

GPT-4: 65.3 (#222), Yi-Lightning: 67.4 (#196)

Instruction Following benchmarks
BenchmarkGPT-4Yi-Lightning
LMArena Instruction Following12411278

Long Context Yi-Lightning leads

GPT-4: 37.7 (#212), Yi-Lightning: 39.4 (#181)

Long Context benchmarks
BenchmarkGPT-4Yi-Lightning
LMArena Longer Query12441297

Writing & Preference Yi-Lightning leads

GPT-4: 34.9 (#268), Yi-Lightning: 50.2 (#184)

Writing & Preference benchmarks
BenchmarkGPT-4Yi-Lightning
LMArena Text12631302
LMArena Creative Writing12441280
LMArena Multi-Turn12571311
EQ-Bench Creative Writing752—

Frequently asked questions

Is GPT-4 better than Yi-Lightning?

Yi-Lightning is the stronger model overall, scoring 37.1 to 29.1 on the Noometry Index.

Is GPT-4 or Yi-Lightning better for coding?

GPT-4 scores higher on coding benchmarks: 31.6 versus 27.4 in the Noometry coding category.

How many benchmarks do GPT-4 and Yi-Lightning share?

17 benchmarks have published results for both models. GPT-4 has 38 scored results on Noometry and Yi-Lightning has 18.

Related comparisons

Go deeper