Model comparison

GPT-4.1 nano vs Llama 4 Maverick

Llama 4 Maverick is the stronger model overall, scoring 30.9 to 27.9 on the Noometry Index. GPT-4.1 nano costs 1.7× less per token, which makes it the better buy when Llama 4 Maverick's lead doesn't matter for your workload.

Last verified . 37 shared benchmarks.

GPT-4.1 nano OpenAI

27.9

Rank #327 Confirmed

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Summary

  • They share 37 benchmarks with published results for both. GPT-4.1 nano scores higher in 2 categories and Llama 4 Maverick in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama 4 Maverick leads 33.4 to 21.8.
  • The biggest single-benchmark swing is MMLU-Pro: 55% for GPT-4.1 nano and 81% for Llama 4 Maverick.
  • GPT-4.1 nano is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.19 / $0.65 for Llama 4 Maverick.
  • GPT-4.1 nano accepts more context: 1.05M tokens versus 128K.
  • Llama 4 Maverick has downloadable open weights; the other is API-only.

Side by side

GPT-4.1 nano and Llama 4 Maverick specifications
GPT-4.1 nanoLlama 4 Maverick
ProviderOpenAIMeta
Noometry Index27.930.9
Released2025-04-142025-04-05
WeightsProprietaryOpen
Context window1.05M128K
Max output33K4K
Input $ / M tokens$0.10$0.19
Output $ / M tokens$0.40$0.65
Results tracked3854

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 4 Maverick leads

GPT-4.1 nano: 24.1 (#330), Llama 4 Maverick: 26.6 (#324)

Coding benchmarks
BenchmarkGPT-4.1 nanoLlama 4 Maverick
Aider Polyglot8.9%15.6%
SciCode25.9%33.1%
WeirdML19%24.5%
LMArena Coding13061302
SWE-bench Verified (bash only)—21%
BigCodeBench Instruct—49.7%
BigCodeBench Complete—61.4%
ALE-Bench—172.97

Agentic & Tool Use Llama 4 Maverick leads

GPT-4.1 nano: 26.5 (#104), Llama 4 Maverick: 28.2 (#91)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1 nanoLlama 4 Maverick
Berkeley Function Calling Leaderboard33%37.3%

Reasoning Llama 4 Maverick leads

GPT-4.1 nano: 8.5 (#349), Llama 4 Maverick: 10.1 (#342)

Reasoning benchmarks
BenchmarkGPT-4.1 nanoLlama 4 Maverick
ARC-AGI-20%0%
Kagi LLM Benchmark33.3%55.9%
ARC-AGI-10%4.4%
CritPt0%0%
LMArena Hard Prompts12861281
DTBench52.5%61.9%
LMCA5.5%15.9%
Epoch Capabilities Index129.62132.2
SimpleBench—27.7%
NYT Connections (extended)—8%
EnigmaEval—0.6%
ForecastBench—57.5

Math Too close to call

GPT-4.1 nano: 26.9 (#252), Llama 4 Maverick: 26.0 (#262)

Math benchmarks
BenchmarkGPT-4.1 nanoLlama 4 Maverick
OTIS Mock AIME 2024-202528.9%20.6%
Omni-MATH36.7%42.2%
LMArena Math12741299
MATH Level 570%73%
FrontierMath (Feb 2025 set)1%0.7%

Knowledge Llama 4 Maverick leads

GPT-4.1 nano: 21.8 (#273), Llama 4 Maverick: 33.4 (#204)

Knowledge benchmarks
BenchmarkGPT-4.1 nanoLlama 4 Maverick
GPQA Diamond48.9%67%
MMLU-Pro55%81%
GPQA (HELM)50.7%65%
LMArena Expert12721259
Humanity's Last Exam—5.7%
SimpleQA Verified6%—
Confabulations—22.6%
Vectara Hallucination Rate—8.2%

Multimodal Llama 4 Maverick leads

GPT-4.1 nano: 29.2 (#113), Llama 4 Maverick: 31.6 (#105)

Multimodal benchmarks
BenchmarkGPT-4.1 nanoLlama 4 Maverick
LMArena Vision10631142
GeoBench—52%
SpatialViz-Bench—31.8%

Multilingual Too close to call

GPT-4.1 nano: 41.6 (#205), Llama 4 Maverick: 42.2 (#195)

Multilingual benchmarks
BenchmarkGPT-4.1 nanoLlama 4 Maverick
LMArena Non-English12601269
LMArena Chinese12701277
LMArena German12881291
LMArena Japanese11981207
LMArena Russian12611286
LMArena French—1259
LMArena Korean—1203
LMArena Spanish—1293

Instruction Following Llama 4 Maverick leads

GPT-4.1 nano: 67.8 (#193), Llama 4 Maverick: 71.7 (#146)

Instruction Following benchmarks
BenchmarkGPT-4.1 nanoLlama 4 Maverick
IFEval84.3%90.8%
LMArena Instruction Following12671267

Long Context Llama 4 Maverick leads

GPT-4.1 nano: 23.7 (#296), Llama 4 Maverick: 31.4 (#279)

Long Context benchmarks
BenchmarkGPT-4.1 nanoLlama 4 Maverick
Fiction.LiveBench25%46.2%
LMArena Longer Query12831280

Writing & Preference GPT-4.1 nano leads

GPT-4.1 nano: 40.5 (#243), Llama 4 Maverick: 38.8 (#252)

Writing & Preference benchmarks
BenchmarkGPT-4.1 nanoLlama 4 Maverick
LMArena Text12851287
LMArena Creative Writing12601267
EQ-Bench Creative Writing946860
WildBench81.2%80%
LMArena Multi-Turn12771289
Short-Story Creative Writing—62%

Frequently asked questions

Is GPT-4.1 nano better than Llama 4 Maverick?

Llama 4 Maverick is the stronger model overall, scoring 30.9 to 27.9 on the Noometry Index. GPT-4.1 nano costs 1.7× less per token, which makes it the better buy when Llama 4 Maverick's lead doesn't matter for your workload.

Which is cheaper, GPT-4.1 nano or Llama 4 Maverick?

GPT-4.1 nano is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; Llama 4 Maverick lists at $0.19 and $0.65.

Is GPT-4.1 nano or Llama 4 Maverick better for coding?

Llama 4 Maverick scores higher on coding benchmarks: 26.6 versus 24.1 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 nano does, with 1.05M tokens against 128K.

How many benchmarks do GPT-4.1 nano and Llama 4 Maverick share?

37 benchmarks have published results for both models. GPT-4.1 nano has 38 scored results on Noometry and Llama 4 Maverick has 54.

Related comparisons

Go deeper