Model comparison

Llama 4 Maverick vs Nova Premier 1.0

Nova Premier 1.0 is the stronger model overall, scoring 38.3 to 30.9 on the Noometry Index. Llama 4 Maverick costs 16× less per token, which makes it the better buy when Nova Premier 1.0's lead doesn't matter for your workload.

Last verified . 6 shared benchmarks.

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Nova Premier 1.0 Amazon

38.3

Rank #189 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Llama 4 Maverick scores higher in 1 category and Nova Premier 1.0 in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Nova Premier 1.0 leads 23.0 to 10.1.
  • The biggest single-benchmark swing is GPQA (HELM): 65% for Llama 4 Maverick and 51.8% for Nova Premier 1.0.
  • Llama 4 Maverick is cheaper at $0.19 / $0.65 per million input/output tokens, against $2.50 / $12.50 for Nova Premier 1.0.
  • Nova Premier 1.0 accepts more context: 1M tokens versus 128K.
  • Llama 4 Maverick has downloadable open weights; the other is API-only.

Side by side

Llama 4 Maverick and Nova Premier 1.0 specifications
Llama 4 MaverickNova Premier 1.0
ProviderMetaAmazon
Noometry Index30.938.3
Released2025-04-052025-04-30
WeightsOpenProprietary
Context window128K1M
Max output4K10K
Input $ / M tokens$0.19$2.50
Output $ / M tokens$0.65$12.50
Results tracked546

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 4 Maverick: 26.6 (#324), Nova Premier 1.0: —

Coding benchmarks
BenchmarkLlama 4 MaverickNova Premier 1.0
SWE-bench Verified (bash only)21%—
Aider Polyglot15.6%—
SciCode33.1%—
WeirdML24.5%—
BigCodeBench Instruct49.7%—
LMArena Coding1302—
BigCodeBench Complete61.4%—
ALE-Bench172.97—

Agentic & Tool Use Not comparable

Llama 4 Maverick: 28.2 (#91), Nova Premier 1.0: —

Agentic & Tool Use benchmarks
BenchmarkLlama 4 MaverickNova Premier 1.0
Berkeley Function Calling Leaderboard37.3%—

Reasoning Nova Premier 1.0 leads

Llama 4 Maverick: 10.1 (#342), Nova Premier 1.0: 23.0 (#185)

Reasoning benchmarks
BenchmarkLlama 4 MaverickNova Premier 1.0
Kagi LLM Benchmark55.9%44.8%
ARC-AGI-20%—
SimpleBench27.7%—
NYT Connections (extended)8%—
ARC-AGI-14.4%—
CritPt0%—
EnigmaEval0.6%—
LMArena Hard Prompts1281—
DTBench61.9%—
LMCA15.9%—
Epoch Capabilities Index132.2—
ForecastBench57.5—

Math Nova Premier 1.0 leads

Llama 4 Maverick: 26.0 (#262), Nova Premier 1.0: 33.8 (#199)

Math benchmarks
BenchmarkLlama 4 MaverickNova Premier 1.0
Omni-MATH42.2%35%
OTIS Mock AIME 2024-202520.6%—
LMArena Math1299—
MATH Level 573%—
FrontierMath (Feb 2025 set)0.7%—

Knowledge Nova Premier 1.0 leads

Llama 4 Maverick: 33.4 (#204), Nova Premier 1.0: 35.8 (#180)

Knowledge benchmarks
BenchmarkLlama 4 MaverickNova Premier 1.0
MMLU-Pro81%72.6%
GPQA (HELM)65%51.8%
GPQA Diamond67%—
Humanity's Last Exam5.7%—
Confabulations22.6%—
Vectara Hallucination Rate8.2%—
LMArena Expert1259—

Multimodal Not comparable

Llama 4 Maverick: 31.6 (#105), Nova Premier 1.0: —

Multimodal benchmarks
BenchmarkLlama 4 MaverickNova Premier 1.0
LMArena Vision1142—
GeoBench52%—
SpatialViz-Bench31.8%—

Multilingual Not comparable

Llama 4 Maverick: 42.2 (#195), Nova Premier 1.0: —

Multilingual benchmarks
BenchmarkLlama 4 MaverickNova Premier 1.0
LMArena Non-English1269—
LMArena Chinese1277—
LMArena French1259—
LMArena German1291—
LMArena Japanese1207—
LMArena Korean1203—
LMArena Russian1286—
LMArena Spanish1293—

Instruction Following Llama 4 Maverick leads

Llama 4 Maverick: 71.7 (#146), Nova Premier 1.0: 66.8 (#204)

Instruction Following benchmarks
BenchmarkLlama 4 MaverickNova Premier 1.0
IFEval90.8%80.3%
LMArena Instruction Following1267—

Long Context Not comparable

Llama 4 Maverick: 31.4 (#279), Nova Premier 1.0: —

Long Context benchmarks
BenchmarkLlama 4 MaverickNova Premier 1.0
Fiction.LiveBench46.2%—
LMArena Longer Query1280—

Writing & Preference Nova Premier 1.0 leads

Llama 4 Maverick: 38.8 (#252), Nova Premier 1.0: 51.0 (#178)

Writing & Preference benchmarks
BenchmarkLlama 4 MaverickNova Premier 1.0
WildBench80%78.8%
LMArena Text1287—
LMArena Creative Writing1267—
Short-Story Creative Writing62%—
EQ-Bench Creative Writing860—
LMArena Multi-Turn1289—

Frequently asked questions

Is Llama 4 Maverick better than Nova Premier 1.0?

Nova Premier 1.0 is the stronger model overall, scoring 38.3 to 30.9 on the Noometry Index. Llama 4 Maverick costs 16× less per token, which makes it the better buy when Nova Premier 1.0's lead doesn't matter for your workload.

Which is cheaper, Llama 4 Maverick or Nova Premier 1.0?

Llama 4 Maverick is cheaper. It lists at $0.19 per million input tokens and $0.65 per million output tokens; Nova Premier 1.0 lists at $2.50 and $12.50.

Which has the bigger context window?

Nova Premier 1.0 does, with 1M tokens against 128K.

How many benchmarks do Llama 4 Maverick and Nova Premier 1.0 share?

6 benchmarks have published results for both models. Llama 4 Maverick has 54 scored results on Noometry and Nova Premier 1.0 has 6.

Related comparisons

Go deeper