Model comparison

Llama 3.1 Tulu 3 8b vs o1-pro

Llama 3.1 Tulu 3 8b is the stronger model overall, scoring 35.7 to 31.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

o1-pro OpenAI

31.5

Rank #271 Reported

Summary

  • Llama 3.1 Tulu 3 8b has downloadable open weights; the other is API-only.

Side by side

Llama 3.1 Tulu 3 8b and o1-pro specifications
Llama 3.1 Tulu 3 8bo1-pro
ProviderAllen Institute for AI (Ai2)OpenAI
Noometry Index35.731.5
Released—2025-03-19
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$150
Output $ / M tokens—$600
Results tracked113

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.1 Tulu 3 8b: 34.4 (#235), o1-pro: —

Coding benchmarks
BenchmarkLlama 3.1 Tulu 3 8bo1-pro
LMArena Coding1183—

Reasoning Llama 3.1 Tulu 3 8b leads

Llama 3.1 Tulu 3 8b: 22.8 (#188), o1-pro: 20.4 (#239)

Reasoning benchmarks
BenchmarkLlama 3.1 Tulu 3 8bo1-pro
ARC-AGI-1—23.3%
EnigmaEval—6.1%
LMArena Hard Prompts1174—

Math Not comparable

Llama 3.1 Tulu 3 8b: 33.9 (#198), o1-pro: —

Math benchmarks
BenchmarkLlama 3.1 Tulu 3 8bo1-pro
LMArena Math1195—

Knowledge Not comparable

Llama 3.1 Tulu 3 8b: —, o1-pro: 29.7 (#234)

Knowledge benchmarks
BenchmarkLlama 3.1 Tulu 3 8bo1-pro
Humanity's Last Exam—8.1%

Multilingual Not comparable

Llama 3.1 Tulu 3 8b: 35.4 (#246), o1-pro: —

Multilingual benchmarks
BenchmarkLlama 3.1 Tulu 3 8bo1-pro
LMArena Non-English1169—
LMArena Chinese1176—
LMArena Russian1193—

Instruction Following Not comparable

Llama 3.1 Tulu 3 8b: 61.3 (#246), o1-pro: —

Instruction Following benchmarks
BenchmarkLlama 3.1 Tulu 3 8bo1-pro
LMArena Instruction Following1174—

Long Context Not comparable

Llama 3.1 Tulu 3 8b: 35.8 (#239), o1-pro: —

Long Context benchmarks
BenchmarkLlama 3.1 Tulu 3 8bo1-pro
LMArena Longer Query1181—

Writing & Preference Not comparable

Llama 3.1 Tulu 3 8b: 39.7 (#245), o1-pro: —

Writing & Preference benchmarks
BenchmarkLlama 3.1 Tulu 3 8bo1-pro
LMArena Text1193—
LMArena Creative Writing1182—
LMArena Multi-Turn1154—

Frequently asked questions

Is Llama 3.1 Tulu 3 8b better than o1-pro?

Llama 3.1 Tulu 3 8b is the stronger model overall, scoring 35.7 to 31.5 on the Noometry Index.

How many benchmarks do Llama 3.1 Tulu 3 8b and o1-pro share?

0 benchmarks have published results for both models. Llama 3.1 Tulu 3 8b has 11 scored results on Noometry and o1-pro has 3.

Related comparisons

Go deeper