Model comparison

Llama 3.2 90B vs o3

o3 is the stronger model overall, scoring 47.5 to 27.5 on the Noometry Index.

Last verified . 7 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Llama 3.2 90B scores higher in 0 categories and o3 in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where o3 leads 50.2 to 11.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 2.6% for Llama 3.2 90B and 84.4% for o3.
  • Llama 3.2 90B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 90B and o3 specifications
Llama 3.2 90Bo3
ProviderMetaOpenAI
Noometry Index27.547.5
Released2024-09-242025-04-16
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked963

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, o3: 46.8 (#64)

Coding benchmarks
BenchmarkLlama 3.2 90Bo3
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
GSO—8.8%
WeirdML—52.4%
LMArena Coding—1408
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use o3 leads

Llama 3.2 90B: 30.0 (#80), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90Bo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
BALROG27.3%—
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Llama 3.2 90B: 21.7 (#217), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkLlama 3.2 90Bo3
EnigmaEval0.4%13.1%
Epoch Capabilities Index125.5146.86
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
LMArena Hard Prompts—1402
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
ForecastBench—62.5

Math o3 leads

Llama 3.2 90B: 11.1 (#308), o3: 50.2 (#58)

Math benchmarks
BenchmarkLlama 3.2 90Bo3
OTIS Mock AIME 2024-20252.6%84.4%
MATH Level 539.4%97.8%
FrontierMath (Tiers 1-3)—33.3%
Omni-MATH—71.4%
LMArena Math—1426
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Llama 3.2 90B: 21.7 (#274), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkLlama 3.2 90Bo3
GPQA Diamond41%81.8%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%
LMArena Expert—1402
MMLU80.3%—

Multimodal o3 leads

Llama 3.2 90B: 25.4 (#124), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkLlama 3.2 90Bo3
LMArena Vision10001214
GeoBench52%74%
VPCT—52%

Multilingual Not comparable

Llama 3.2 90B: —, o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkLlama 3.2 90Bo3
LMArena Non-English—1401
LMArena Chinese—1437
LMArena French—1430
LMArena German—1420
LMArena Japanese—1403
LMArena Korean—1370
LMArena Russian—1406
LMArena Spanish—1395

Instruction Following Not comparable

Llama 3.2 90B: —, o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkLlama 3.2 90Bo3
IFEval—86.9%
LMArena Instruction Following—1368

Long Context Not comparable

Llama 3.2 90B: —, o3: 53.3 (#6)

Long Context benchmarks
BenchmarkLlama 3.2 90Bo3
Fiction.LiveBench—88.9%
CL-bench—17.8%
LMArena Longer Query—1372

Writing & Preference Not comparable

Llama 3.2 90B: —, o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90Bo3
LMArena Text—1410
LMArena Creative Writing—1359
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%
LMArena Multi-Turn—1405

Frequently asked questions

Is Llama 3.2 90B better than o3?

o3 is the stronger model overall, scoring 47.5 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and o3 share?

7 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper