Model comparison
Llama 3.1-405B vs o1-pro
Llama 3.1-405B and o1-pro score almost the same on the Noometry Index (30.7 vs 31.5), so choose on price, context window or the category you care about most.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in reasoning, where o1-pro leads 20.4 to 16.8.
- Llama 3.1-405B has downloadable open weights; the other is API-only.
Side by side
| Llama 3.1-405B | o1-pro | |
|---|---|---|
| Provider | Meta | OpenAI |
| Noometry Index | 30.7 | 31.5 |
| Released | 2024-07-23 | 2025-03-19 |
| Weights | Open | Proprietary |
| Context window | — | 200K |
| Max output | — | 100K |
| Input $ / M tokens | — | $150 |
| Output $ / M tokens | — | $600 |
| Results tracked | 42 | 3 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Llama 3.1-405B: 33.1 (#262), o1-pro: —
| Benchmark | Llama 3.1-405B | o1-pro |
|---|---|---|
| WeirdML | 21.4% | — |
| LMArena Coding | 1291 | — |
Agentic & Tool Use Not comparable
Llama 3.1-405B: 21.0 (#140), o1-pro: —
| Benchmark | Llama 3.1-405B | o1-pro |
|---|---|---|
| TheAgentCompany | 7.4% | — |
| Cybench | 7.5% | — |
Reasoning o1-pro leads
Llama 3.1-405B: 16.8 (#300), o1-pro: 20.4 (#239)
| Benchmark | Llama 3.1-405B | o1-pro |
|---|---|---|
| SimpleBench | 23% | — |
| Kagi LLM Benchmark | 45% | — |
| ARC-AGI-1 | — | 23.3% |
| EnigmaEval | — | 6.1% |
| LMArena Hard Prompts | 1269 | — |
| DTBench | 61.4% | — |
| BIG-Bench Hard | 82.9% | — |
| Epoch Capabilities Index | 128.75 | — |
| ForecastBench | 59.9 | — |
| HellaSwag | 89.2% | — |
| PIQA | 85.9% | — |
| WinoGrande | 89.2% | — |
Math Not comparable
Llama 3.1-405B: 18.4 (#290), o1-pro: —
| Benchmark | Llama 3.1-405B | o1-pro |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 9.7% | — |
| Omni-MATH | 24.9% | — |
| LMArena Math | 1281 | — |
| MATH Level 5 | 49.8% | — |
Knowledge Too close to call
Llama 3.1-405B: 30.4 (#227), o1-pro: 29.7 (#234)
| Benchmark | Llama 3.1-405B | o1-pro |
|---|---|---|
| GPQA Diamond | 50.9% | — |
| Humanity's Last Exam | — | 8.1% |
| MMLU-Pro | 72.3% | — |
| Confabulations | 17.6% | — |
| GPQA (HELM) | 52.2% | — |
| LMArena Expert | 1243 | — |
| ARC (AI2) Challenge | 95.3% | — |
| MMLU | 84.5% | — |
| TriviaQA | 82.7% | — |
Multilingual Not comparable
Llama 3.1-405B: 40.7 (#214), o1-pro: —
| Benchmark | Llama 3.1-405B | o1-pro |
|---|---|---|
| LMArena Non-English | 1248 | — |
| LMArena Chinese | 1242 | — |
| LMArena French | 1279 | — |
| LMArena German | 1252 | — |
| LMArena Japanese | 1208 | — |
| LMArena Korean | 1184 | — |
| LMArena Russian | 1265 | — |
| LMArena Spanish | 1260 | — |
Instruction Following Not comparable
Llama 3.1-405B: 65.9 (#214), o1-pro: —
| Benchmark | Llama 3.1-405B | o1-pro |
|---|---|---|
| IFEval | 81.1% | — |
| LMArena Instruction Following | 1259 | — |
Long Context Not comparable
Llama 3.1-405B: 38.4 (#197), o1-pro: —
| Benchmark | Llama 3.1-405B | o1-pro |
|---|---|---|
| LMArena Longer Query | 1266 | — |
Writing & Preference Not comparable
Llama 3.1-405B: 38.9 (#251), o1-pro: —
| Benchmark | Llama 3.1-405B | o1-pro |
|---|---|---|
| LMArena Text | 1284 | — |
| LMArena Creative Writing | 1262 | — |
| EQ-Bench Creative Writing | 870 | — |
| WildBench | 78.3% | — |
| LMArena Multi-Turn | 1297 | — |
Frequently asked questions
Is Llama 3.1-405B better than o1-pro?
Llama 3.1-405B and o1-pro score almost the same on the Noometry Index (30.7 vs 31.5), so choose on price, context window or the category you care about most.
How many benchmarks do Llama 3.1-405B and o1-pro share?
0 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and o1-pro has 3.