Model comparison

Llama 2-70B vs o3

o3 is the stronger model overall, scoring 47.5 to 24.4 on the Noometry Index.

Last verified . 23 shared benchmarks.

Llama 2-70B Meta

24.4

Rank #349 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Llama 2-70B scores higher in 0 categories and o3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 7.4.
  • The biggest single-benchmark swing is MATH Level 5: 3.3% for Llama 2-70B and 97.8% for o3.
  • Llama 2-70B has downloadable open weights; the other is API-only.

Side by side

Llama 2-70B and o3 specifications
Llama 2-70Bo3
ProviderMetaOpenAI
Noometry Index24.447.5
Released2023-07-182025-04-16
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked3563

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Llama 2-70B: 31.4 (#286), o3: 46.8 (#64)

Coding benchmarks
BenchmarkLlama 2-70Bo3
LMArena Coding10791408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Llama 2-70B: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkLlama 2-70Bo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Llama 2-70B: 14.4 (#325), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkLlama 2-70Bo3
LMArena Hard Prompts10731402
DTBench41.6%84.8%
Epoch Capabilities Index113.79146.86
ForecastBench51.462.5
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
LMCA—39.7%
BIG-Bench Hard64.9%—
CommonsenseQA 2.050%—
HellaSwag85.3%—
LAMBADA78.9%—
PIQA82.8%—
WinoGrande80.2%—

Math o3 leads

Llama 2-70B: 8.1 (#326), o3: 50.2 (#58)

Math benchmarks
BenchmarkLlama 2-70Bo3
OTIS Mock AIME 2024-20250%84.4%
LMArena Math10911426
MATH Level 53.3%97.8%
FrontierMath (Tiers 1-3)—33.3%
Omni-MATH—71.4%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%
GSM8K69.6%—

Knowledge o3 leads

Llama 2-70B: 7.4 (#310), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkLlama 2-70Bo3
GPQA Diamond26.3%81.8%
LMArena Expert10391402
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%
ARC (AI2) Challenge78.3%—
BoolQ88.6%—
MMLU69.9%—
OpenBookQA60.2%—
TriviaQA87.6%—

Multimodal Not comparable

Llama 2-70B: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkLlama 2-70Bo3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual o3 leads

Llama 2-70B: 27.7 (#274), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkLlama 2-70Bo3
LMArena Non-English10451401
LMArena Chinese9951437
LMArena French10901430
LMArena German10411420
LMArena Japanese9271403
LMArena Korean9641370
LMArena Russian10831406
LMArena Spanish11431395

Instruction Following o3 leads

Llama 2-70B: 54.9 (#278), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkLlama 2-70Bo3
LMArena Instruction Following10711368
IFEval—86.9%

Long Context o3 leads

Llama 2-70B: 32.3 (#270), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkLlama 2-70Bo3
LMArena Longer Query10621372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Llama 2-70B: 32.3 (#279), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkLlama 2-70Bo3
LMArena Text11151410
LMArena Creative Writing10751359
LMArena Multi-Turn10881405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Llama 2-70B better than o3?

o3 is the stronger model overall, scoring 47.5 to 24.4 on the Noometry Index.

Is Llama 2-70B or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 31.4 in the Noometry coding category.

How many benchmarks do Llama 2-70B and o3 share?

23 benchmarks have published results for both models. Llama 2-70B has 35 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper