Model comparison

GPT-4 vs Llama-3.3-70B-Instruct

Llama-3.3-70B-Instruct is the stronger model overall, scoring 30.6 to 29.1 on the Noometry Index.

Last verified . 28 shared benchmarks.

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 28 benchmarks with published results for both. GPT-4 scores higher in 4 categories and Llama-3.3-70B-Instruct in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama-3.3-70B-Instruct leads 47.6 to 34.9.
  • The biggest single-benchmark swing is MATH Level 5: 23% for GPT-4 and 41.6% for Llama-3.3-70B-Instruct.
  • Llama-3.3-70B-Instruct is cheaper at $0.10 / $0.32 per million input/output tokens, against $30 / $60 for GPT-4.
  • Llama-3.3-70B-Instruct accepts more context: 128K tokens versus 8K.
  • Llama-3.3-70B-Instruct has downloadable open weights; the other is API-only.

Side by side

GPT-4 and Llama-3.3-70B-Instruct specifications
GPT-4Llama-3.3-70B-Instruct
ProviderOpenAIMeta
Noometry Index29.130.6
Released2023-03-142024-12-06
WeightsProprietaryOpen
Context window8K128K
Max output8K4K
Input $ / M tokens$30$0.10
Output $ / M tokens$60$0.32
Results tracked3843

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-4: 31.6 (#283), Llama-3.3-70B-Instruct: 31.0 (#290)

Coding benchmarks
BenchmarkGPT-4Llama-3.3-70B-Instruct
WeirdML12.4%14.4%
BigCodeBench Instruct46%46.9%
LMArena Coding12541268
BigCodeBench Complete57.2%57.5%
SciCode—26%
LiveBench Coding—36.6%
HumanEval+79.3%—

Agentic & Tool Use Not comparable

GPT-4: —, Llama-3.3-70B-Instruct: 25.8 (#105)

Agentic & Tool Use benchmarks
BenchmarkGPT-4Llama-3.3-70B-Instruct
Berkeley Function Calling Leaderboard—31.9%
BALROG—23%
METR Time Horizons36.1%—

Reasoning GPT-4 leads

GPT-4: 17.8 (#289), Llama-3.3-70B-Instruct: 14.1 (#327)

Reasoning benchmarks
BenchmarkGPT-4Llama-3.3-70B-Instruct
LMArena Hard Prompts12411257
DTBench62.7%59.5%
LMCA17.1%17.5%
Epoch Capabilities Index125.89127.33
ForecastBench57.858.6
SimpleBench—19.9%
CritPt—0%
Chess Puzzles4%—
LiveBench Reasoning—50.8%
Mystery Game Puzzles12%—
LiveBench Data Analysis—49.5%
BIG-Bench Hard75.1%—
HellaSwag95.3%—
LiveBench—50.2%
WinoGrande87.5%—

Math Llama-3.3-70B-Instruct leads

GPT-4: 10.8 (#309), Llama-3.3-70B-Instruct: 15.3 (#298)

Math benchmarks
BenchmarkGPT-4Llama-3.3-70B-Instruct
OTIS Mock AIME 2024-20251.1%5.1%
LMArena Math12691267
MATH Level 523%41.6%
LiveBench Math—42.2%
GSM8K92%—

Knowledge Llama-3.3-70B-Instruct leads

GPT-4: 18.4 (#282), Llama-3.3-70B-Instruct: 30.6 (#226)

Knowledge benchmarks
BenchmarkGPT-4Llama-3.3-70B-Instruct
GPQA Diamond35.7%47.4%
LMArena Expert12111225
MMLU86.4%86.3%
Confabulations—22.8%
Vectara Hallucination Rate—4.1%
TriviaQA84.8%—

Multilingual Too close to call

GPT-4: 40.6 (#215), Llama-3.3-70B-Instruct: 39.9 (#220)

Multilingual benchmarks
BenchmarkGPT-4Llama-3.3-70B-Instruct
LMArena Non-English12461236
LMArena Chinese12421217
LMArena French12831281
LMArena German12511251
LMArena Japanese12091150
LMArena Korean11841143
LMArena Russian12511252
LMArena Spanish12611270

Instruction Following Llama-3.3-70B-Instruct leads

GPT-4: 65.3 (#222), Llama-3.3-70B-Instruct: 71.1 (#157)

Instruction Following benchmarks
BenchmarkGPT-4Llama-3.3-70B-Instruct
LMArena Instruction Following12411242
LiveBench Instruction Following—82.7%

Long Context GPT-4 leads

GPT-4: 37.7 (#212), Llama-3.3-70B-Instruct: 26.4 (#295)

Long Context benchmarks
BenchmarkGPT-4Llama-3.3-70B-Instruct
LMArena Longer Query12441256
Fiction.LiveBench—33.3%

Writing & Preference Llama-3.3-70B-Instruct leads

GPT-4: 34.9 (#268), Llama-3.3-70B-Instruct: 47.6 (#207)

Writing & Preference benchmarks
BenchmarkGPT-4Llama-3.3-70B-Instruct
LMArena Text12631274
LMArena Creative Writing12441250
LMArena Multi-Turn12571280
EQ-Bench Creative Writing752—
LiveBench Language—39.2%

Frequently asked questions

Is GPT-4 better than Llama-3.3-70B-Instruct?

Llama-3.3-70B-Instruct is the stronger model overall, scoring 30.6 to 29.1 on the Noometry Index.

Which is cheaper, GPT-4 or Llama-3.3-70B-Instruct?

Llama-3.3-70B-Instruct is cheaper. It lists at $0.10 per million input tokens and $0.32 per million output tokens; GPT-4 lists at $30 and $60.

Is GPT-4 or Llama-3.3-70B-Instruct better for coding?

They score almost the same on coding (31.6 vs 31.0); test both on your own repository before choosing.

Which has the bigger context window?

Llama-3.3-70B-Instruct does, with 128K tokens against 8K.

How many benchmarks do GPT-4 and Llama-3.3-70B-Instruct share?

28 benchmarks have published results for both models. GPT-4 has 38 scored results on Noometry and Llama-3.3-70B-Instruct has 43.

Related comparisons

Go deeper