Model comparison

GPT-4 vs Llama 4 Maverick

Llama 4 Maverick is the stronger model overall, scoring 30.9 to 29.1 on the Noometry Index.

Last verified . 28 shared benchmarks.

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Summary

  • They share 28 benchmarks with published results for both. GPT-4 scores higher in 3 categories and Llama 4 Maverick in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama 4 Maverick leads 26.0 to 10.8.
  • The biggest single-benchmark swing is MATH Level 5: 23% for GPT-4 and 73% for Llama 4 Maverick.
  • Llama 4 Maverick is cheaper at $0.19 / $0.65 per million input/output tokens, against $30 / $60 for GPT-4.
  • Llama 4 Maverick accepts more context: 128K tokens versus 8K.
  • Llama 4 Maverick has downloadable open weights; the other is API-only.

Side by side

GPT-4 and Llama 4 Maverick specifications
GPT-4Llama 4 Maverick
ProviderOpenAIMeta
Noometry Index29.130.9
Released2023-03-142025-04-05
WeightsProprietaryOpen
Context window8K128K
Max output8K4K
Input $ / M tokens$30$0.19
Output $ / M tokens$60$0.65
Results tracked3854

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4 leads

GPT-4: 31.6 (#283), Llama 4 Maverick: 26.6 (#324)

Coding benchmarks
BenchmarkGPT-4Llama 4 Maverick
WeirdML12.4%24.5%
BigCodeBench Instruct46%49.7%
LMArena Coding12541302
BigCodeBench Complete57.2%61.4%
SWE-bench Verified (bash only)—21%
Aider Polyglot—15.6%
SciCode—33.1%
ALE-Bench—172.97
HumanEval+79.3%—

Agentic & Tool Use Not comparable

GPT-4: —, Llama 4 Maverick: 28.2 (#91)

Agentic & Tool Use benchmarks
BenchmarkGPT-4Llama 4 Maverick
Berkeley Function Calling Leaderboard—37.3%
METR Time Horizons36.1%—

Reasoning GPT-4 leads

GPT-4: 17.8 (#289), Llama 4 Maverick: 10.1 (#342)

Reasoning benchmarks
BenchmarkGPT-4Llama 4 Maverick
LMArena Hard Prompts12411281
DTBench62.7%61.9%
LMCA17.1%15.9%
Epoch Capabilities Index125.89132.2
ForecastBench57.857.5
ARC-AGI-2—0%
SimpleBench—27.7%
Kagi LLM Benchmark—55.9%
NYT Connections (extended)—8%
ARC-AGI-1—4.4%
CritPt—0%
Chess Puzzles4%—
EnigmaEval—0.6%
Mystery Game Puzzles12%—
BIG-Bench Hard75.1%—
HellaSwag95.3%—
WinoGrande87.5%—

Math Llama 4 Maverick leads

GPT-4: 10.8 (#309), Llama 4 Maverick: 26.0 (#262)

Math benchmarks
BenchmarkGPT-4Llama 4 Maverick
OTIS Mock AIME 2024-20251.1%20.6%
LMArena Math12691299
MATH Level 523%73%
Omni-MATH—42.2%
FrontierMath (Feb 2025 set)—0.7%
GSM8K92%—

Knowledge Llama 4 Maverick leads

GPT-4: 18.4 (#282), Llama 4 Maverick: 33.4 (#204)

Knowledge benchmarks
BenchmarkGPT-4Llama 4 Maverick
GPQA Diamond35.7%67%
LMArena Expert12111259
Humanity's Last Exam—5.7%
MMLU-Pro—81%
Confabulations—22.6%
Vectara Hallucination Rate—8.2%
GPQA (HELM)—65%
MMLU86.4%—
TriviaQA84.8%—

Multimodal Not comparable

GPT-4: —, Llama 4 Maverick: 31.6 (#105)

Multimodal benchmarks
BenchmarkGPT-4Llama 4 Maverick
LMArena Vision—1142
GeoBench—52%
SpatialViz-Bench—31.8%

Multilingual Llama 4 Maverick leads

GPT-4: 40.6 (#215), Llama 4 Maverick: 42.2 (#195)

Multilingual benchmarks
BenchmarkGPT-4Llama 4 Maverick
LMArena Non-English12461269
LMArena Chinese12421277
LMArena French12831259
LMArena German12511291
LMArena Japanese12091207
LMArena Korean11841203
LMArena Russian12511286
LMArena Spanish12611293

Instruction Following Llama 4 Maverick leads

GPT-4: 65.3 (#222), Llama 4 Maverick: 71.7 (#146)

Instruction Following benchmarks
BenchmarkGPT-4Llama 4 Maverick
LMArena Instruction Following12411267
IFEval—90.8%

Long Context GPT-4 leads

GPT-4: 37.7 (#212), Llama 4 Maverick: 31.4 (#279)

Long Context benchmarks
BenchmarkGPT-4Llama 4 Maverick
LMArena Longer Query12441280
Fiction.LiveBench—46.2%

Writing & Preference Llama 4 Maverick leads

GPT-4: 34.9 (#268), Llama 4 Maverick: 38.8 (#252)

Writing & Preference benchmarks
BenchmarkGPT-4Llama 4 Maverick
LMArena Text12631287
LMArena Creative Writing12441267
EQ-Bench Creative Writing752860
LMArena Multi-Turn12571289
Short-Story Creative Writing—62%
WildBench—80%

Frequently asked questions

Is GPT-4 better than Llama 4 Maverick?

Llama 4 Maverick is the stronger model overall, scoring 30.9 to 29.1 on the Noometry Index.

Which is cheaper, GPT-4 or Llama 4 Maverick?

Llama 4 Maverick is cheaper. It lists at $0.19 per million input tokens and $0.65 per million output tokens; GPT-4 lists at $30 and $60.

Is GPT-4 or Llama 4 Maverick better for coding?

GPT-4 scores higher on coding benchmarks: 31.6 versus 26.6 in the Noometry coding category.

Which has the bigger context window?

Llama 4 Maverick does, with 128K tokens against 8K.

How many benchmarks do GPT-4 and Llama 4 Maverick share?

28 benchmarks have published results for both models. GPT-4 has 38 scored results on Noometry and Llama 4 Maverick has 54.

Related comparisons

Go deeper