Model comparison

Claude Opus 4.8 vs Llama 4 Maverick

Claude Opus 4.8 is the stronger model overall, scoring 60.7 to 30.9 on the Noometry Index. Llama 4 Maverick costs 33× less per token, which makes it the better buy when Claude Opus 4.8's lead doesn't matter for your workload.

Last verified . 36 shared benchmarks.

Claude Opus 4.8 Anthropic

60.7

Rank #13 Confirmed

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Summary

  • They share 36 benchmarks with published results for both. Claude Opus 4.8 scores higher in 10 categories and Llama 4 Maverick in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Opus 4.8 leads 64.7 to 10.1.
  • The biggest single-benchmark swing is ARC-AGI-1: 92.5% for Claude Opus 4.8 and 4.4% for Llama 4 Maverick.
  • Llama 4 Maverick is cheaper at $0.19 / $0.65 per million input/output tokens, against $5 / $25 for Claude Opus 4.8.
  • Claude Opus 4.8 accepts more context: 1M tokens versus 128K.
  • Llama 4 Maverick has downloadable open weights; the other is API-only.

Side by side

Claude Opus 4.8 and Llama 4 Maverick specifications
Claude Opus 4.8Llama 4 Maverick
ProviderAnthropicMeta
Noometry Index60.730.9
Released2026-05-282025-04-05
WeightsProprietaryOpen
Context window1M128K
Max output128K4K
Input $ / M tokens$5$0.19
Output $ / M tokens$25$0.65
Results tracked6554

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.8 leads

Claude Opus 4.8: 59.9 (#12), Llama 4 Maverick: 26.6 (#324)

Coding benchmarks
BenchmarkClaude Opus 4.8Llama 4 Maverick
SciCode53.5%33.1%
WeirdML82.9%24.5%
LMArena Coding14901302
ALE-Bench1,564172.97
DeepSWE59%—
FrontierCode46.5%—
SWE-bench Verified (bash only)—21%
Aider Polyglot—15.6%
LMArena WebDev1556—
GSO47.1%—
BigCodeBench Instruct—49.7%
BigCodeBench Complete—61.4%

Agentic & Tool Use Claude Opus 4.8 leads

Claude Opus 4.8: 47.6 (#11), Llama 4 Maverick: 28.2 (#91)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.8Llama 4 Maverick
APEX-Agents48.9%—
Berkeley Function Calling Leaderboard—37.3%
OSWorld 2.020.6%—
Remote Labor Index8.3%—
τ²-bench Banking39.7%—
DeepResearch Bench50.2%—
PostTrainBench33.8%—
GBAEval70.9%—
GDP.pdf24%—
LMArena Search1204—
Vending-Bench 25,787—

Reasoning Claude Opus 4.8 leads

Claude Opus 4.8: 64.7 (#16), Llama 4 Maverick: 10.1 (#342)

Reasoning benchmarks
BenchmarkClaude Opus 4.8Llama 4 Maverick
ARC-AGI-272.1%0%
SimpleBench64.8%27.7%
Kagi LLM Benchmark88.8%55.9%
NYT Connections (extended)91.1%8%
ARC-AGI-192.5%4.4%
CritPt20.9%0%
EnigmaEval23.5%0.6%
LMArena Hard Prompts14821281
DTBench94.9%61.9%
LMCA57.5%15.9%
Epoch Capabilities Index158.21132.2
ForecastBench59.957.5
Chess Puzzles34%—
EBR-Bench28.6%—
Mystery Game Puzzles36%—
Surface Evolver Bench87.5%—
Bench to the Future 30.14—

Math Claude Opus 4.8 leads

Claude Opus 4.8: 78.4 (#13), Llama 4 Maverick: 26.0 (#262)

Math benchmarks
BenchmarkClaude Opus 4.8Llama 4 Maverick
OTIS Mock AIME 2024-202598.3%20.6%
LMArena Math14871299
FrontierMath (Feb 2025 set)47.2%0.7%
FrontierMath (Tiers 1-3)80%—
FrontierMath Tier 456.1%—
MathArena Final-Answer Competitions91.8%—
ProofBench69%—
Omni-MATH—42.2%
MATH Level 5—73%
FrontierMath Tier 4 (v1)31.3%—

Knowledge Claude Opus 4.8 leads

Claude Opus 4.8: 61.3 (#29), Llama 4 Maverick: 33.4 (#204)

Knowledge benchmarks
BenchmarkClaude Opus 4.8Llama 4 Maverick
GPQA Diamond91%67%
LMArena Expert15021259
Humanity's Last Exam—5.7%
SimpleQA Verified53%—
MMLU-Pro—81%
Confabulations—22.6%
Vectara Hallucination Rate—8.2%
GPQA (HELM)—65%

Multimodal Claude Opus 4.8 leads

Claude Opus 4.8: 42.9 (#26), Llama 4 Maverick: 31.6 (#105)

Multimodal benchmarks
BenchmarkClaude Opus 4.8Llama 4 Maverick
LMArena Vision12941142
GeoBench—52%
Blueprint-Bench 214.5%—
Furniture Assembly42.5%—
LMArena Document1475—
SpatialViz-Bench—31.8%

Multilingual Claude Opus 4.8 leads

Claude Opus 4.8: 55.2 (#33), Llama 4 Maverick: 42.2 (#195)

Multilingual benchmarks
BenchmarkClaude Opus 4.8Llama 4 Maverick
LMArena Non-English14501269
LMArena Chinese15071277
LMArena French14811259
LMArena German14721291
LMArena Japanese14401207
LMArena Korean14321203
LMArena Russian14741286
LMArena Spanish14661293

Instruction Following Claude Opus 4.8 leads

Claude Opus 4.8: 77.4 (#24), Llama 4 Maverick: 71.7 (#146)

Instruction Following benchmarks
BenchmarkClaude Opus 4.8Llama 4 Maverick
LMArena Instruction Following14761267
IFEval—90.8%

Long Context Claude Opus 4.8 leads

Claude Opus 4.8: 45.4 (#35), Llama 4 Maverick: 31.4 (#279)

Long Context benchmarks
BenchmarkClaude Opus 4.8Llama 4 Maverick
LMArena Longer Query14831280
Fiction.LiveBench—46.2%

Writing & Preference Claude Opus 4.8 leads

Claude Opus 4.8: 72.0 (#16), Llama 4 Maverick: 38.8 (#252)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.8Llama 4 Maverick
LMArena Text14611287
LMArena Creative Writing14541267
EQ-Bench Creative Writing1840860
LMArena Multi-Turn14761289
Short-Story Creative Writing—62%
WildBench—80%
EQ-Bench 41281—

Frequently asked questions

Is Claude Opus 4.8 better than Llama 4 Maverick?

Claude Opus 4.8 is the stronger model overall, scoring 60.7 to 30.9 on the Noometry Index. Llama 4 Maverick costs 33× less per token, which makes it the better buy when Claude Opus 4.8's lead doesn't matter for your workload.

Which is cheaper, Claude Opus 4.8 or Llama 4 Maverick?

Llama 4 Maverick is cheaper. It lists at $0.19 per million input tokens and $0.65 per million output tokens; Claude Opus 4.8 lists at $5 and $25.

Is Claude Opus 4.8 or Llama 4 Maverick better for coding?

Claude Opus 4.8 scores higher on coding benchmarks: 59.9 versus 26.6 in the Noometry coding category.

Which has the bigger context window?

Claude Opus 4.8 does, with 1M tokens against 128K.

How many benchmarks do Claude Opus 4.8 and Llama 4 Maverick share?

36 benchmarks have published results for both models. Claude Opus 4.8 has 65 scored results on Noometry and Llama 4 Maverick has 54.

Related comparisons

Go deeper