Meta, open weights
Llama 4 Scout
Llama 4 Scout by Meta ranks 330th of 354 ranked models on the Noometry Index as of October 2026, with a score of 27.7. Its strongest category is multimodal, where it ranks 102nd. API pricing starts at $0.10 per million input tokens and $0.30 per million output tokens, with a 128K-token context window.
Last verified
Specifications
- Noometry rank
- #330 of 354
- Index score
- 27.7
- Evidence
- Confirmed 43 results
- Provider
Meta
- Released
- April 5, 2025
- Weights
- Open weights
- Reasoning
- No
- Context window
- 128K
- Max output
- 4K
- Input price
- $0.10 / M
- Output price
- $0.30 / M
- Blended price
- $0.15 / M
- Output speed
- 272 tokens/s Kagi
- Value
- #45 of 219
- Knowledge cutoff
- August 2024
- Input
- text, image
- Hugging Face
- meta-llama/Llama-4-Scout-17B-16E-Instruct
Category scores
Each category score combines every public result we have in that category.
- Coding 20.2
- Agentic & Tool Use 24.6
- Reasoning 9.1
- Math 19.6
- Knowledge 31.9
- Multimodal 32.2
- Multilingual 41.0
- Instruction Following 65.8
- Long Context 27.5
- Writing & Preference 37.0
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 20.2 | #339 | 4 |
| Agentic & Tool Use | 24.6 | #119 | 1 |
| Reasoning | 9.1 | #345 | 7 |
| Math | 19.6 | #286 | 4 |
| Knowledge | 31.9 | #217 | 5 |
| Multimodal | 32.2 | #102 | 1 |
| Multilingual | 41.0 | #212 | 1 |
| Instruction Following | 65.8 | #217 | 2 |
| Long Context | 27.5 | #294 | 2 |
| Writing & Preference | 37.0 | #261 | 5 |
Strengths and weaknesses
Categories where Llama 4 Scout places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Knowledge | 31.9 | −5.4 | #217 of 314, top 70% |
| Instruction Following | 65.8 | −5.5 | #217 of 305, top 72% |
| Multilingual | 41.0 | −6.4 | #212 of 297, top 72% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Coding | 20.2 | −18.5 | #339 of 340, top 100% |
| Long Context | 27.5 | −13.4 | #294 of 296, top 100% |
| Reasoning | 9.1 | −14.5 | #345 of 350, top 99% |
Closest competitors
The models ranked just above and below Llama 4 Scout. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Gemma 3 4B | #326 | 28.1 | $0.05 | 72 | Compare |
| GPT-4.1 nano | #327 | 27.9 | $0.18 | 135 | Compare |
| Phi 3 Mini 4k Instruct | #328 | 27.9 | — | — | Compare |
| Yi-34B | #329 | 27.8 | — | — | Compare |
| Llama 3.2 90B | #331 | 27.5 | — | — | Compare |
| Gemini 1.0 Pro | #332 | 27.3 | — | — | Compare |
| Mixtral 8x22B | #333 | 27.1 | $3 | — | Compare |
| Mixtral 8x7B | #334 | 27.1 | $0.70 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified (bash only) | 9.1% | #38 of 39, top 98% | SWE-bench | 2025-07-20 | |
| SciCode | 17% | #119 of 121, top 99% | Epoch AI | ||
| LMArena Coding | 1286 | #209 of 294, top 72% | LMArena | 2026-10-08 | |
| BigCodeBench Complete | 43.1% | #47 of 66, top 72% | BigCodeBench | 2025-04-05 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 28.1% | #37 of 49, top 76% | fc | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 0% | #81 of 83, top 98% | Epoch AI | ||
| Kagi LLM Benchmark | 36.9% | #86 of 99, top 87% | Kagi LLM Benchmark | ||
| ARC-AGI-1 | 0.5% | #82 of 83, top 99% | Epoch AI | ||
| CritPt | 0% | #120 of 134, top 90% | Epoch AI | ||
| LMArena Hard Prompts | 1266 | #211 of 297, top 72% | LMArena | 2026-10-08 | |
| DTBench | 57.9% | #120 of 151, top 80% | Epoch AI | ||
| LMCA | 12% | #109 of 125, top 88% | Epoch AI | ||
| Epoch Capabilities Index | 129.64 | #139 of 213, top 66% | Epoch AI | 2025-04-05 | |
| ForecastBench | 57.5 | #57 of 72, top 80% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 7.8% | #138 of 173, top 80% | Epoch AI | 2025-04-08 | |
| Omni-MATH | 37.3% | #30 of 57, top 53% | HELM Capabilities | ||
| LMArena Math | 1287 | #188 of 285, top 66% | LMArena | 2026-10-08 | |
| MATH Level 5 | 62.3% | #38 of 79, top 49% | Epoch AI | 2025-04-08 | |
| FrontierMath (Feb 2025 set) | 0% | #68 of 68, top 100% | Epoch AI | 2025-04-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 51.8% | #125 of 186, top 68% | Epoch AI | 2025-04-08 | |
| MMLU-Pro | 74.2% | #27 of 58, top 47% | HELM Capabilities | ||
| Vectara Hallucination Rate (lower is better) | 7.7% | #34 of 96, top 36% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 50.7% | #34 of 57, top 60% | HELM Capabilities | ||
| LMArena Expert | 1235 | #205 of 273, top 76% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1118 | #105 of 122, top 87% | LMArena | 2026-10-09 | |
| SpatialViz-Bench | 34.2% | #4 of 8, top 50% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1252 | #212 of 297, top 72% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1255 | #204 of 285, top 72% | LMArena | 2026-10-08 | |
| LMArena French | 1282 | #169 of 223, top 76% | LMArena | 2026-10-08 | |
| LMArena German | 1272 | #165 of 231, top 72% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1206 | #157 of 211, top 75% | LMArena | 2026-10-08 | |
| LMArena Korean | 1207 | #156 of 213, top 74% | LMArena | 2026-10-08 | |
| LMArena Russian | 1263 | #202 of 283, top 72% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1278 | #170 of 226, top 76% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 81.8% | #34 of 57, top 60% | HELM Capabilities | ||
| LMArena Instruction Following | 1248 | #215 of 298, top 73% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 36% | #44 of 47, top 94% | Epoch AI | ||
| LMArena Longer Query | 1265 | #213 of 291, top 74% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1279 | #210 of 297, top 71% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1249 | #206 of 295, top 70% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 783 | #105 of 115, top 92% | EQ-Bench | ||
| WildBench | 78% | #41 of 57, top 72% | HELM Capabilities | ||
| LMArena Multi-Turn | 1280 | #200 of 295, top 68% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $0.20 | $0.78 | — | 2026-10-10 |
| bedrock | $0.17 | $0.66 | — | 2026-10-10 |
| deepinfra | $0.10 | $0.30 | — | 2026-10-10 |
| openrouter | $0.10 | $0.30 | — | 2026-10-10 |
Compare Llama 4 Scout
- Llama 4 Scout vs Yi-34B
- Llama 4 Scout vs Llama 3.2 90B
- Llama 4 Scout vs Phi 3 Mini 4k Instruct
- Llama 4 Scout vs Gemini 1.0 Pro
- Llama 4 Scout vs GPT-4.1 nano
- Llama 4 Scout vs Mixtral 8x22B
- Llama 4 Scout vs GPT-6 Astra
- Llama 4 Scout vs Claude Fable 5.1
- Llama 4 Scout vs Gemini 3.8 Flash
- Llama 4 Scout vs Kimi K3
- Llama 4 Scout vs Grok 4.6
- Llama 4 Scout vs Qwen3.8 Max
- Llama 4 Scout vs GLM-5.3
- Llama 4 Scout vs DeepSeek V4 Pro
Other Meta models
- Muse Spark 1.354.8
- Muse Spark50.6
- Muse Spark 1.250.3
- Muse Spark 1.149.9
- Muse Glimmer41.7
- Codellama 70b Instruct33.7
- Llama 4 Maverick30.9
- Codellama 34b Instruct30.8
Frequently asked questions
How good is Llama 4 Scout?
Llama 4 Scout by Meta ranks 330th of 354 ranked models on the Noometry Index as of October 2026, with a score of 27.7. Its strongest category is multimodal, where it ranks 102nd. API pricing starts at $0.10 per million input tokens and $0.30 per million output tokens, with a 128K-token context window.
How much does Llama 4 Scout cost?
Llama 4 Scout costs $0.10 per million input tokens and $0.30 per million output tokens on deepinfra.
What is Llama 4 Scout's context window?
Llama 4 Scout accepts up to 128K tokens of input and can write up to 4K tokens in one response.
Is Llama 4 Scout open source?
Yes. Llama 4 Scout's weights are downloadable from Hugging Face (meta-llama/Llama-4-Scout-17B-16E-Instruct); check the license for commercial terms.
How fast is Llama 4 Scout?
Llama 4 Scout generated about 272 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Llama 4 Scout's strengths and weaknesses?
Relative to other ranked models, Llama 4 Scout places best in knowledge, instruction following, multilingual and lowest in coding, long context, reasoning.
What is Llama 4 Scout best at?
Its best category is multimodal, where it ranks 102nd on Noometry.