Model comparison

GPT-4.5 vs Llama-3.3-70B-Instruct

GPT-4.5 is the stronger model overall, scoring 37.2 to 30.6 on the Noometry Index.

Last verified . 32 shared benchmarks.

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 32 benchmarks with published results for both. GPT-4.5 scores higher in 8 categories and Llama-3.3-70B-Instruct in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-4.5 leads 32.6 to 15.3.
  • The biggest single-benchmark swing is LiveBench Coding: 75.2% for GPT-4.5 and 36.6% for Llama-3.3-70B-Instruct.
  • Llama-3.3-70B-Instruct has downloadable open weights; the other is API-only.

Side by side

GPT-4.5 and Llama-3.3-70B-Instruct specifications
GPT-4.5Llama-3.3-70B-Instruct
ProviderOpenAIMeta
Noometry Index37.230.6
Released2025-02-272024-12-06
WeightsProprietaryOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.32
Results tracked4243

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4.5 leads

GPT-4.5: 42.2 (#109), Llama-3.3-70B-Instruct: 31.0 (#290)

Coding benchmarks
BenchmarkGPT-4.5Llama-3.3-70B-Instruct
WeirdML39.4%14.4%
LiveBench Coding75.2%36.6%
LMArena Coding13961268
Aider Polyglot44.9%—
SciCode—26%
BigCodeBench Instruct—46.9%
BigCodeBench Complete—57.5%

Agentic & Tool Use GPT-4.5 leads

GPT-4.5: 27.9 (#97), Llama-3.3-70B-Instruct: 25.8 (#105)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.5Llama-3.3-70B-Instruct
Berkeley Function Calling Leaderboard—31.9%
Cybench17.5%—
BALROG—23%

Reasoning Too close to call

GPT-4.5: 13.9 (#330), Llama-3.3-70B-Instruct: 14.1 (#327)

Reasoning benchmarks
BenchmarkGPT-4.5Llama-3.3-70B-Instruct
SimpleBench34.5%19.9%
LiveBench Reasoning71.1%50.8%
LMArena Hard Prompts14031257
LiveBench Data Analysis64.3%49.5%
Epoch Capabilities Index136.74127.33
ForecastBench61.758.6
LiveBench69%50.2%
ARC-AGI-20.8%—
ARC-AGI-110.3%—
CritPt—0%
EnigmaEval3.2%—
DTBench—59.5%
LMCA—17.5%

Math GPT-4.5 leads

GPT-4.5: 32.6 (#211), Llama-3.3-70B-Instruct: 15.3 (#298)

Math benchmarks
BenchmarkGPT-4.5Llama-3.3-70B-Instruct
OTIS Mock AIME 2024-202537.8%5.1%
LiveBench Math69.3%42.2%
LMArena Math14121267
MATH Level 578.6%41.6%

Knowledge GPT-4.5 leads

GPT-4.5: 32.5 (#211), Llama-3.3-70B-Instruct: 30.6 (#226)

Knowledge benchmarks
BenchmarkGPT-4.5Llama-3.3-70B-Instruct
GPQA Diamond68.7%47.4%
Confabulations13.6%22.8%
LMArena Expert13941225
Humanity's Last Exam5.4%—
Vectara Hallucination Rate—4.1%
MMLU—86.3%

Multimodal Not comparable

GPT-4.5: 37.6 (#71), Llama-3.3-70B-Instruct: —

Multimodal benchmarks
BenchmarkGPT-4.5Llama-3.3-70B-Instruct
LMArena Vision1195—
VPCT45%—

Multilingual GPT-4.5 leads

GPT-4.5: 52.5 (#83), Llama-3.3-70B-Instruct: 39.9 (#220)

Multilingual benchmarks
BenchmarkGPT-4.5Llama-3.3-70B-Instruct
LMArena Non-English14131236
LMArena Chinese14211217
LMArena French14181281
LMArena German14571251
LMArena Japanese14161150
LMArena Korean13921143
LMArena Russian14191252
LMArena Spanish—1270

Instruction Following GPT-4.5 leads

GPT-4.5: 72.6 (#134), Llama-3.3-70B-Instruct: 71.1 (#157)

Instruction Following benchmarks
BenchmarkGPT-4.5Llama-3.3-70B-Instruct
LiveBench Instruction Following72.3%82.7%
LMArena Instruction Following14041242

Long Context GPT-4.5 leads

GPT-4.5: 40.4 (#155), Llama-3.3-70B-Instruct: 26.4 (#295)

Long Context benchmarks
BenchmarkGPT-4.5Llama-3.3-70B-Instruct
Fiction.LiveBench63.9%33.3%
LMArena Longer Query14061256

Writing & Preference GPT-4.5 leads

GPT-4.5: 56.9 (#134), Llama-3.3-70B-Instruct: 47.6 (#207)

Writing & Preference benchmarks
BenchmarkGPT-4.5Llama-3.3-70B-Instruct
LMArena Text14171274
LMArena Creative Writing13941250
LMArena Multi-Turn14441280
LiveBench Language61.5%39.2%
Short-Story Creative Writing75.6%—
EQ-Bench Creative Writing1258—

Frequently asked questions

Is GPT-4.5 better than Llama-3.3-70B-Instruct?

GPT-4.5 is the stronger model overall, scoring 37.2 to 30.6 on the Noometry Index.

Is GPT-4.5 or Llama-3.3-70B-Instruct better for coding?

GPT-4.5 scores higher on coding benchmarks: 42.2 versus 31.0 in the Noometry coding category.

How many benchmarks do GPT-4.5 and Llama-3.3-70B-Instruct share?

32 benchmarks have published results for both models. GPT-4.5 has 42 scored results on Noometry and Llama-3.3-70B-Instruct has 43.

Related comparisons

Go deeper