Model comparison

Llama 3.1-405B vs o1-pro

Llama 3.1-405B and o1-pro score almost the same on the Noometry Index (30.7 vs 31.5), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

o1-pro OpenAI

31.5

Rank #271 Reported

Summary

  • The widest gap is in reasoning, where o1-pro leads 20.4 to 16.8.
  • Llama 3.1-405B has downloadable open weights; the other is API-only.

Side by side

Llama 3.1-405B and o1-pro specifications
Llama 3.1-405Bo1-pro
ProviderMetaOpenAI
Noometry Index30.731.5
Released2024-07-232025-03-19
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$150
Output $ / M tokens—$600
Results tracked423

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.1-405B: 33.1 (#262), o1-pro: —

Coding benchmarks
BenchmarkLlama 3.1-405Bo1-pro
WeirdML21.4%—
LMArena Coding1291—

Agentic & Tool Use Not comparable

Llama 3.1-405B: 21.0 (#140), o1-pro: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405Bo1-pro
TheAgentCompany7.4%—
Cybench7.5%—

Reasoning o1-pro leads

Llama 3.1-405B: 16.8 (#300), o1-pro: 20.4 (#239)

Reasoning benchmarks
BenchmarkLlama 3.1-405Bo1-pro
SimpleBench23%—
Kagi LLM Benchmark45%—
ARC-AGI-1—23.3%
EnigmaEval—6.1%
LMArena Hard Prompts1269—
DTBench61.4%—
BIG-Bench Hard82.9%—
Epoch Capabilities Index128.75—
ForecastBench59.9—
HellaSwag89.2%—
PIQA85.9%—
WinoGrande89.2%—

Math Not comparable

Llama 3.1-405B: 18.4 (#290), o1-pro: —

Math benchmarks
BenchmarkLlama 3.1-405Bo1-pro
OTIS Mock AIME 2024-20259.7%—
Omni-MATH24.9%—
LMArena Math1281—
MATH Level 549.8%—

Knowledge Too close to call

Llama 3.1-405B: 30.4 (#227), o1-pro: 29.7 (#234)

Knowledge benchmarks
BenchmarkLlama 3.1-405Bo1-pro
GPQA Diamond50.9%—
Humanity's Last Exam—8.1%
MMLU-Pro72.3%—
Confabulations17.6%—
GPQA (HELM)52.2%—
LMArena Expert1243—
ARC (AI2) Challenge95.3%—
MMLU84.5%—
TriviaQA82.7%—

Multilingual Not comparable

Llama 3.1-405B: 40.7 (#214), o1-pro: —

Multilingual benchmarks
BenchmarkLlama 3.1-405Bo1-pro
LMArena Non-English1248—
LMArena Chinese1242—
LMArena French1279—
LMArena German1252—
LMArena Japanese1208—
LMArena Korean1184—
LMArena Russian1265—
LMArena Spanish1260—

Instruction Following Not comparable

Llama 3.1-405B: 65.9 (#214), o1-pro: —

Instruction Following benchmarks
BenchmarkLlama 3.1-405Bo1-pro
IFEval81.1%—
LMArena Instruction Following1259—

Long Context Not comparable

Llama 3.1-405B: 38.4 (#197), o1-pro: —

Long Context benchmarks
BenchmarkLlama 3.1-405Bo1-pro
LMArena Longer Query1266—

Writing & Preference Not comparable

Llama 3.1-405B: 38.9 (#251), o1-pro: —

Writing & Preference benchmarks
BenchmarkLlama 3.1-405Bo1-pro
LMArena Text1284—
LMArena Creative Writing1262—
EQ-Bench Creative Writing870—
WildBench78.3%—
LMArena Multi-Turn1297—

Frequently asked questions

Is Llama 3.1-405B better than o1-pro?

Llama 3.1-405B and o1-pro score almost the same on the Noometry Index (30.7 vs 31.5), so choose on price, context window or the category you care about most.

How many benchmarks do Llama 3.1-405B and o1-pro share?

0 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and o1-pro has 3.

Related comparisons

Go deeper