Model comparison

Llama 3.2 90B vs o1

o1 is the stronger model overall, scoring 40.9 to 27.5 on the Noometry Index.

Last verified . 7 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Llama 3.2 90B scores higher in 1 category and o1 in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where o1 leads 36.1 to 11.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 2.6% for Llama 3.2 90B and 73.3% for o1.
  • Llama 3.2 90B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 90B and o1 specifications
Llama 3.2 90Bo1
ProviderMetaOpenAI
Noometry Index27.540.9
Released2024-09-242024-09-12
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked952

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, o1: 46.1 (#70)

Coding benchmarks
BenchmarkLlama 3.2 90Bo1
Aider Polyglot—61.7%
WeirdML—47.6%
LiveBench Coding—69.7%
LMArena Coding—1367
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Llama 3.2 90B leads

Llama 3.2 90B: 30.0 (#80), o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90Bo1
Cybench—10%
BALROG27.3%—
METR Time Horizons—51.1%

Reasoning o1 leads

Llama 3.2 90B: 21.7 (#217), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkLlama 3.2 90Bo1
EnigmaEval0.4%5.7%
Epoch Capabilities Index125.5141.91
SimpleBench—41.7%
ARC-AGI-1—30.7%
Chess Puzzles—15%
LiveBench Reasoning—91.6%
LMArena Hard Prompts—1371
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
LiveBench—75.7%

Math o1 leads

Llama 3.2 90B: 11.1 (#308), o1: 36.1 (#175)

Math benchmarks
BenchmarkLlama 3.2 90Bo1
OTIS Mock AIME 2024-20252.6%73.3%
MATH Level 539.4%94.7%
FrontierMath (Tiers 1-3)—14.7%
LiveBench Math—80.3%
LMArena Math—1388
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

Llama 3.2 90B: 21.7 (#274), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkLlama 3.2 90Bo1
GPQA Diamond41%76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%
LMArena Expert—1361
MMLU80.3%—

Multimodal o1 leads

Llama 3.2 90B: 25.4 (#124), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkLlama 3.2 90Bo1
LMArena Vision10001168
GeoBench52%80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual Not comparable

Llama 3.2 90B: —, o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkLlama 3.2 90Bo1
LMArena Non-English—1358
LMArena Chinese—1394
LMArena French—1344
LMArena German—1337
LMArena Japanese—1346
LMArena Korean—1396
LMArena Russian—1356
LMArena Spanish—1345

Instruction Following Not comparable

Llama 3.2 90B: —, o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkLlama 3.2 90Bo1
LiveBench Instruction Following—81.5%
LMArena Instruction Following—1367

Long Context Not comparable

Llama 3.2 90B: —, o1: 50.3 (#9)

Long Context benchmarks
BenchmarkLlama 3.2 90Bo1
Fiction.LiveBench—83.3%
LMArena Longer Query—1378

Writing & Preference Not comparable

Llama 3.2 90B: —, o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90Bo1
LMArena Text—1366
LMArena Creative Writing—1348
Short-Story Creative Writing—70.2%
LMArena Multi-Turn—1369
LiveBench Language—65.4%

Frequently asked questions

Is Llama 3.2 90B better than o1?

o1 is the stronger model overall, scoring 40.9 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and o1 share?

7 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper