Model comparison

GPT-4 vs Llama 3-8B

GPT-4 is the stronger model overall, scoring 29.1 to 25.5 on the Noometry Index.

Last verified . 30 shared benchmarks.

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Llama 3-8B Meta

25.5

Rank #344 Confirmed

Summary

  • They share 30 benchmarks with published results for both. GPT-4 scores higher in 7 categories and Llama 3-8B in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-4 leads 18.4 to 7.8.
  • The biggest single-benchmark swing is BigCodeBench Complete: 57.2% for GPT-4 and 36.9% for Llama 3-8B.
  • Llama 3-8B has downloadable open weights; the other is API-only.

Side by side

GPT-4 and Llama 3-8B specifications
GPT-4Llama 3-8B
ProviderOpenAIMeta
Noometry Index29.125.5
Released2023-03-142024-04-18
WeightsProprietaryOpen
Context window8K—
Max output8K—
Input $ / M tokens$30—
Output $ / M tokens$60—
Results tracked3834

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-4: 31.6 (#283), Llama 3-8B: 31.0 (#289)

Coding benchmarks
BenchmarkGPT-4Llama 3-8B
BigCodeBench Instruct46%31.9%
LMArena Coding12541152
BigCodeBench Complete57.2%36.9%
HumanEval+79.3%56.7%
WeirdML12.4%—
MBPP+—54.8%

Agentic & Tool Use Not comparable

GPT-4: —, Llama 3-8B: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4Llama 3-8B
METR Time Horizons36.1%—

Reasoning GPT-4 leads

GPT-4: 17.8 (#289), Llama 3-8B: 14.3 (#326)

Reasoning benchmarks
BenchmarkGPT-4Llama 3-8B
Chess Puzzles4%0%
LMArena Hard Prompts12411133
DTBench62.7%43.9%
Epoch Capabilities Index125.89116.45
ForecastBench57.858.6
WinoGrande87.5%75.7%
Mystery Game Puzzles12%—
LMCA17.1%—
Adversarial NLI—57.3%
BIG-Bench Hard75.1%—
HellaSwag95.3%—

Math GPT-4 leads

GPT-4: 10.8 (#309), Llama 3-8B: 8.8 (#323)

Math benchmarks
BenchmarkGPT-4Llama 3-8B
OTIS Mock AIME 2024-20251.1%1.9%
LMArena Math12691151
MATH Level 523%6.1%
GSM8K92%—

Knowledge GPT-4 leads

GPT-4: 18.4 (#282), Llama 3-8B: 7.8 (#308)

Knowledge benchmarks
BenchmarkGPT-4Llama 3-8B
GPQA Diamond35.7%26.1%
LMArena Expert12111113
MMLU86.4%68.8%
TriviaQA84.8%67.7%
ARC (AI2) Challenge—82.8%
OpenBookQA—82.6%

Multilingual GPT-4 leads

GPT-4: 40.6 (#215), Llama 3-8B: 30.8 (#261)

Multilingual benchmarks
BenchmarkGPT-4Llama 3-8B
LMArena Non-English12461098
LMArena Chinese12421076
LMArena French12831159
LMArena German12511104
LMArena Japanese1209967
LMArena Korean11841004
LMArena Russian12511109
LMArena Spanish12611173

Instruction Following GPT-4 leads

GPT-4: 65.3 (#222), Llama 3-8B: 58.4 (#260)

Instruction Following benchmarks
BenchmarkGPT-4Llama 3-8B
LMArena Instruction Following12411127

Long Context GPT-4 leads

GPT-4: 37.7 (#212), Llama 3-8B: 34.2 (#251)

Long Context benchmarks
BenchmarkGPT-4Llama 3-8B
LMArena Longer Query12441128

Writing & Preference Llama 3-8B leads

GPT-4: 34.9 (#268), Llama 3-8B: 37.5 (#256)

Writing & Preference benchmarks
BenchmarkGPT-4Llama 3-8B
LMArena Text12631166
LMArena Creative Writing12441150
LMArena Multi-Turn12571152
EQ-Bench Creative Writing752—

Frequently asked questions

Is GPT-4 better than Llama 3-8B?

GPT-4 is the stronger model overall, scoring 29.1 to 25.5 on the Noometry Index.

Is GPT-4 or Llama 3-8B better for coding?

They score almost the same on coding (31.6 vs 31.0); test both on your own repository before choosing.

How many benchmarks do GPT-4 and Llama 3-8B share?

30 benchmarks have published results for both models. GPT-4 has 38 scored results on Noometry and Llama 3-8B has 34.

Related comparisons

Go deeper