Meta, open weights
Llama 3.1-405B
Llama 3.1-405B by Meta ranks 288th of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.7. Its strongest category is agentic & tool use, where it ranks 140th.
Last verified
Specifications
- Noometry rank
- #288 of 354
- Index score
- 30.7
- Evidence
- Confirmed 42 results
- Provider
Meta
- Released
- July 23, 2024
- Weights
- Open weights
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- 78 tokens/s Kagi
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 33.1
- Agentic & Tool Use 21.0
- Reasoning 16.8
- Math 18.4
- Knowledge 30.4
- Multilingual 40.7
- Instruction Following 65.9
- Long Context 38.4
- Writing & Preference 38.9
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 33.1 | #262 | 2 |
| Agentic & Tool Use | 21.0 | #140 | 2 |
| Reasoning | 16.8 | #300 | 4 |
| Math | 18.4 | #290 | 4 |
| Knowledge | 30.4 | #227 | 5 |
| Multilingual | 40.7 | #214 | 1 |
| Instruction Following | 65.9 | #214 | 2 |
| Long Context | 38.4 | #197 | 1 |
| Writing & Preference | 38.9 | #251 | 5 |
Strengths and weaknesses
Categories where Llama 3.1-405B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 38.4 | −2.5 | #197 of 296, top 67% |
| Instruction Following | 65.9 | −5.4 | #214 of 305, top 71% |
| Multilingual | 40.7 | −6.7 | #214 of 297, top 73% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 21.0 | −9.3 | #140 of 154, top 91% |
| Math | 18.4 | −18.2 | #290 of 327, top 89% |
| Reasoning | 16.8 | −6.8 | #300 of 350, top 86% |
Closest competitors
The models ranked just above and below Llama 3.1-405B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Gemma 3 27B | #284 | 30.8 | $0.10 | 62 | Compare |
| Qwen1.5-72B | #285 | 30.8 | — | — | Compare |
| Granite 3.0 2b Instruct | #286 | 30.8 | — | — | Compare |
| Codellama 34b Instruct | #287 | 30.8 | — | — | Compare |
| Yi-1.5-34B | #289 | 30.6 | — | — | Compare |
| Codestral | #290 | 30.6 | $0.45 | 271 | Compare |
| Llama-3.3-70B-Instruct | #291 | 30.6 | $0.16 | — | Compare |
| GPT-4 Turbo | #292 | 30.5 | $15 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| WeirdML | 21.4% | #104 of 119, top 88% | Epoch AI | ||
| LMArena Coding | 1283 | LMArena | 2026-10-08 | ||
| LMArena Coding | 1291 | #204 of 294, top 70% | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| TheAgentCompany | 7.4% | #9 of 14, top 65% | Epoch AI | ||
| Cybench | 7.5% | #19 of 21, top 91% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SimpleBench | 23% | #68 of 77, top 89% | Epoch AI | ||
| Kagi LLM Benchmark | 45% | #70 of 99, top 71% | Kagi LLM Benchmark | ||
| LMArena Hard Prompts | 1263 | LMArena | 2026-10-08 | ||
| LMArena Hard Prompts | 1269 | #207 of 297, top 70% | LMArena | 2026-10-08 | |
| DTBench | 61.4% | #113 of 151, top 75% | Epoch AI | ||
| BIG-Bench Hard | 82.9% | #3 of 27, top 12% | Epoch AI | ||
| Epoch Capabilities Index | 128.75 | #144 of 213, top 68% | Epoch AI | 2024-07-23 | |
| ForecastBench | 59.9 | #35 of 72, top 49% | Epoch AI | ||
| HellaSwag | 89.2% | #3 of 29, top 11% | Epoch AI | ||
| PIQA | 85.9% | #4 of 27, top 15% | Epoch AI | ||
| WinoGrande | 89.2% | Best of 43 | Epoch AI | ||
| WinoGrande | 82.2% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 9.7% | #133 of 173, top 77% | Epoch AI | 2025-02-25 | |
| Omni-MATH | 24.9% | #44 of 57, top 78% | HELM Capabilities | ||
| LMArena Math | 1281 | #192 of 285, top 68% | LMArena | 2026-10-08 | |
| LMArena Math | 1278 | LMArena | 2026-10-08 | ||
| MATH Level 5 | 49.8% | #46 of 79, top 59% | Epoch AI | 2025-01-27 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 50.9% | #127 of 186, top 69% | Epoch AI | 2025-01-27 | |
| MMLU-Pro | 72.3% | #33 of 58, top 57% | HELM Capabilities | ||
| Confabulations (lower is better) | 17.6% | #27 of 51, top 53% | Lech Mazur benchmarks | ||
| GPQA (HELM) | 52.2% | #30 of 57, top 53% | HELM Capabilities | ||
| LMArena Expert | 1243 | #202 of 273, top 74% | LMArena | 2026-10-08 | |
| LMArena Expert | 1229 | LMArena | 2026-10-08 | ||
| ARC (AI2) Challenge | 95.3% | #2 of 39, top 6% | Epoch AI | ||
| MMLU | 84.5% | #10 of 81, top 13% | Epoch AI | ||
| MMLU | 84.4% | Epoch AI | |||
| TriviaQA | 82.7% | #8 of 25, top 32% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1248 | #214 of 297, top 73% | LMArena | 2026-10-08 | |
| LMArena Non-English | 1247 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1242 | #213 of 285, top 75% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1234 | LMArena | 2026-10-08 | ||
| LMArena French | 1271 | LMArena | 2026-10-08 | ||
| LMArena French | 1279 | #172 of 223, top 78% | LMArena | 2026-10-08 | |
| LMArena German | 1252 | #176 of 231, top 77% | LMArena | 2026-10-08 | |
| LMArena German | 1251 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1171 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1208 | #153 of 211, top 73% | LMArena | 2026-10-08 | |
| LMArena Korean | 1172 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1184 | #171 of 213, top 81% | LMArena | 2026-10-08 | |
| LMArena Russian | 1265 | #199 of 283, top 71% | LMArena | 2026-10-08 | |
| LMArena Russian | 1256 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1260 | #179 of 226, top 80% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1253 | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 81.1% | #38 of 57, top 67% | HELM Capabilities | ||
| LMArena Instruction Following | 1259 | #202 of 298, top 68% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1259 | #202 of 298, top 68% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1260 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1266 | #211 of 291, top 73% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1284 | #205 of 297, top 70% | LMArena | 2026-10-08 | |
| LMArena Text | 1282 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1260 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1262 | #197 of 295, top 67% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 870 | #101 of 115, top 88% | EQ-Bench | ||
| WildBench | 78.3% | #40 of 57, top 71% | HELM Capabilities | ||
| LMArena Multi-Turn | 1297 | #189 of 295, top 65% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1287 | LMArena | 2026-10-08 |
Compare Llama 3.1-405B
- Llama 3.1-405B vs Llama 3-70B
- Llama 3.1-405B vs Codellama 34b Instruct
- Llama 3.1-405B vs Yi-1.5-34B
- Llama 3.1-405B vs Granite 3.0 2b Instruct
- Llama 3.1-405B vs Codestral
- Llama 3.1-405B vs Qwen1.5-72B
- Llama 3.1-405B vs Llama-3.3-70B-Instruct
- Llama 3.1-405B vs GPT-6 Astra
- Llama 3.1-405B vs Claude Fable 5.1
- Llama 3.1-405B vs Gemini 3.8 Flash
- Llama 3.1-405B vs Kimi K3
- Llama 3.1-405B vs Grok 4.6
- Llama 3.1-405B vs Qwen3.8 Max
- Llama 3.1-405B vs GLM-5.3
Other Meta models
- Muse Spark 1.354.8
- Muse Spark50.6
- Muse Spark 1.250.3
- Muse Spark 1.149.9
- Muse Glimmer41.7
- Codellama 70b Instruct33.7
- Llama 4 Maverick30.9
- Codellama 34b Instruct30.8
Frequently asked questions
How good is Llama 3.1-405B?
Llama 3.1-405B by Meta ranks 288th of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.7. Its strongest category is agentic & tool use, where it ranks 140th.
Is Llama 3.1-405B open source?
Yes. Llama 3.1-405B's weights are downloadable; check the license for commercial terms.
How fast is Llama 3.1-405B?
Llama 3.1-405B generated about 78 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Llama 3.1-405B's strengths and weaknesses?
Relative to other ranked models, Llama 3.1-405B places best in long context, instruction following, multilingual and lowest in agentic & tool use, math, reasoning.
What is Llama 3.1-405B best at?
Its best category is agentic & tool use, where it ranks 140th on Noometry.