Meta, open weights
Llama 4 Maverick
Llama 4 Maverick by Meta ranks 282nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.9. Its strongest category is agentic & tool use, where it ranks 91st. API pricing starts at $0.19 per million input tokens and $0.65 per million output tokens, with a 128K-token context window.
Last verified
Specifications
- Noometry rank
- #282 of 354
- Index score
- 30.9
- Evidence
- Confirmed 54 results
- Provider
Meta
- Released
- April 5, 2025
- Weights
- Open weights
- Reasoning
- No
- Context window
- 128K
- Max output
- 4K
- Input price
- $0.19 / M
- Output price
- $0.65 / M
- Blended price
- $0.30 / M
- Output speed
- 456 tokens/s Kagi
- Value
- #58 of 219
- Knowledge cutoff
- August 2024
- Input
- text, image
- Hugging Face
- meta-llama/Llama-4-Maverick-17B-128E-Instruct
Category scores
Each category score combines every public result we have in that category.
- Coding 26.6
- Agentic & Tool Use 28.2
- Reasoning 10.1
- Math 26.0
- Knowledge 33.4
- Multimodal 31.6
- Multilingual 42.2
- Instruction Following 71.7
- Long Context 31.4
- Writing & Preference 38.8
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 26.6 | #324 | 7 |
| Agentic & Tool Use | 28.2 | #91 | 1 |
| Reasoning | 10.1 | #342 | 10 |
| Math | 26.0 | #262 | 4 |
| Knowledge | 33.4 | #204 | 7 |
| Multimodal | 31.6 | #105 | 2 |
| Multilingual | 42.2 | #195 | 1 |
| Instruction Following | 71.7 | #146 | 2 |
| Long Context | 31.4 | #279 | 2 |
| Writing & Preference | 38.8 | #252 | 6 |
Strengths and weaknesses
Categories where Llama 4 Maverick places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 71.7 | +0.4 | #146 of 305, top 48% |
| Agentic & Tool Use | 28.2 | −2.2 | #91 of 154, top 60% |
| Knowledge | 33.4 | −3.9 | #204 of 314, top 65% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 10.1 | −13.5 | #342 of 350, top 98% |
| Coding | 26.6 | −12.2 | #324 of 340, top 96% |
| Long Context | 31.4 | −9.6 | #279 of 296, top 95% |
Closest competitors
The models ranked just above and below Llama 4 Maverick. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Mistral Small 3 | #278 | 31.2 | $0.0575 | — | Compare |
| Phi-4 | #279 | 31.2 | $0.0875 | — | Compare |
| Mistral Small 3.2 | #280 | 31.2 | $0.13 | 68 | Compare |
| Amazon Nova Pro | #281 | 31.0 | $1.40 | — | Compare |
| Phi-4 Mini | #283 | 30.9 | $0.13 | — | Compare |
| Gemma 3 27B | #284 | 30.8 | $0.10 | 62 | Compare |
| Qwen1.5-72B | #285 | 30.8 | — | — | Compare |
| Granite 3.0 2b Instruct | #286 | 30.8 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified (bash only) | 21% | #36 of 39, top 93% | SWE-bench | 2025-07-20 | |
| Aider Polyglot | 15.6% | #38 of 44, top 87% | Epoch AI | ||
| SciCode | 33.1% | #103 of 121, top 86% | Epoch AI | ||
| WeirdML | 24.5% | #101 of 119, top 85% | Epoch AI | ||
| BigCodeBench Instruct | 49.7% | #3 of 64, top 5% | BigCodeBench | 2025-04-05 | |
| LMArena Coding | 1302 | #198 of 294, top 68% | LMArena | 2026-10-08 | |
| BigCodeBench Complete | 61.4% | #2 of 66, top 4% | BigCodeBench | 2025-04-05 | |
| ALE-Bench | 172.97 | #104 of 105, top 100% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 37.3% | #26 of 49, top 54% | fc | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 0% | #80 of 83, top 97% | Epoch AI | ||
| SimpleBench | 27.7% | #61 of 77, top 80% | Epoch AI | ||
| Kagi LLM Benchmark | 55.9% | #49 of 99, top 50% | Kagi LLM Benchmark | ||
| NYT Connections (extended) | 8% | #89 of 91, top 98% | Lech Mazur benchmarks | ||
| ARC-AGI-1 | 4.4% | #80 of 83, top 97% | Epoch AI | ||
| CritPt | 0% | #118 of 134, top 89% | Epoch AI | ||
| CritPt | 0% | #118 of 134, top 89% | Epoch AI | ||
| EnigmaEval | 0.6% | #37 of 38, top 98% | Epoch AI | ||
| LMArena Hard Prompts | 1281 | #200 of 297, top 68% | LMArena | 2026-10-08 | |
| DTBench | 61.9% | #110 of 151, top 73% | Epoch AI | ||
| LMCA | 15.9% | #103 of 125, top 83% | Epoch AI | ||
| Epoch Capabilities Index | 132.2 | #133 of 213, top 63% | Epoch AI | 2025-04-06 | |
| ForecastBench | 57.5 | #56 of 72, top 78% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 20.6% | #126 of 173, top 73% | Epoch AI | 2025-04-08 | |
| Omni-MATH | 42.2% | #24 of 57, top 43% | HELM Capabilities | ||
| LMArena Math | 1299 | #185 of 285, top 65% | LMArena | 2026-10-08 | |
| MATH Level 5 | 73% | #29 of 79, top 37% | Epoch AI | 2025-04-08 | |
| FrontierMath (Feb 2025 set) | 0.7% | #62 of 68, top 92% | Epoch AI | 2025-04-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 67% | #102 of 186, top 55% | Epoch AI | 2025-04-08 | |
| Humanity's Last Exam | 5.7% | #33 of 41, top 81% | Epoch AI | ||
| MMLU-Pro | 81% | #15 of 58, top 26% | HELM Capabilities | ||
| Confabulations (lower is better) | 22.6% | #38 of 51, top 75% | Lech Mazur benchmarks | ||
| Vectara Hallucination Rate (lower is better) | 8.2% | #37 of 96, top 39% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 65% | #19 of 57, top 34% | HELM Capabilities | ||
| LMArena Expert | 1259 | #190 of 273, top 70% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1142 | #99 of 122, top 82% | LMArena | 2026-10-09 | |
| GeoBench | 52% | #20 of 25, top 80% | Epoch AI | ||
| SpatialViz-Bench | 31.8% | #8 of 8, top 100% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1269 | #195 of 297, top 66% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1277 | #194 of 285, top 69% | LMArena | 2026-10-08 | |
| LMArena French | 1259 | #181 of 223, top 82% | LMArena | 2026-10-08 | |
| LMArena German | 1291 | #152 of 231, top 66% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1207 | #155 of 211, top 74% | LMArena | 2026-10-08 | |
| LMArena Korean | 1203 | #159 of 213, top 75% | LMArena | 2026-10-08 | |
| LMArena Russian | 1286 | #184 of 283, top 66% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1293 | #162 of 226, top 72% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 90.8% | #8 of 57, top 15% | HELM Capabilities | ||
| LMArena Instruction Following | 1267 | #198 of 298, top 67% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 46.2% | #38 of 47, top 81% | Epoch AI | ||
| LMArena Longer Query | 1280 | #203 of 291, top 70% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1287 | #201 of 297, top 68% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1267 | #194 of 295, top 66% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 62% | #37 of 39, top 95% | Epoch AI | ||
| EQ-Bench Creative Writing | 860 | #102 of 115, top 89% | EQ-Bench | ||
| WildBench | 80% | #31 of 57, top 55% | HELM Capabilities | ||
| LMArena Multi-Turn | 1289 | #196 of 295, top 67% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $0.25 | $1 | — | 2026-10-10 |
| bedrock | $0.24 | $0.97 | — | 2026-10-10 |
| deepinfra | $0.20 | $0.80 | — | 2026-10-10 |
| openrouter | $0.19 | $0.65 | $0.05 | 2026-10-10 |
| vertex | $0.35 | $1.15 | — | 2026-10-10 |
Compare Llama 4 Maverick
- Llama 4 Maverick vs Amazon Nova Pro
- Llama 4 Maverick vs Phi-4 Mini
- Llama 4 Maverick vs Mistral Small 3.2
- Llama 4 Maverick vs Gemma 3 27B
- Llama 4 Maverick vs Phi-4
- Llama 4 Maverick vs Qwen1.5-72B
- Llama 4 Maverick vs GPT-6 Astra
- Llama 4 Maverick vs Claude Fable 5.1
- Llama 4 Maverick vs Gemini 3.8 Flash
- Llama 4 Maverick vs Kimi K3
- Llama 4 Maverick vs Grok 4.6
- Llama 4 Maverick vs Qwen3.8 Max
- Llama 4 Maverick vs GLM-5.3
- Llama 4 Maverick vs DeepSeek V4 Pro
Other Meta models
- Muse Spark 1.354.8
- Muse Spark50.6
- Muse Spark 1.250.3
- Muse Spark 1.149.9
- Muse Glimmer41.7
- Codellama 70b Instruct33.7
- Codellama 34b Instruct30.8
- Llama 3.1-405B30.7
Frequently asked questions
How good is Llama 4 Maverick?
Llama 4 Maverick by Meta ranks 282nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.9. Its strongest category is agentic & tool use, where it ranks 91st. API pricing starts at $0.19 per million input tokens and $0.65 per million output tokens, with a 128K-token context window.
How much does Llama 4 Maverick cost?
Llama 4 Maverick costs $0.19 per million input tokens and $0.65 per million output tokens on openrouter, with cached input at $0.05.
What is Llama 4 Maverick's context window?
Llama 4 Maverick accepts up to 128K tokens of input and can write up to 4K tokens in one response.
Is Llama 4 Maverick open source?
Yes. Llama 4 Maverick's weights are downloadable from Hugging Face (meta-llama/Llama-4-Maverick-17B-128E-Instruct); check the license for commercial terms.
How fast is Llama 4 Maverick?
Llama 4 Maverick generated about 456 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Llama 4 Maverick's strengths and weaknesses?
Relative to other ranked models, Llama 4 Maverick places best in instruction following, agentic & tool use, knowledge and lowest in reasoning, coding, long context.
What is Llama 4 Maverick best at?
Its best category is agentic & tool use, where it ranks 91st on Noometry.