Model comparison

Llama-3.3-70B-Instruct vs o1-mini

o1-mini is the stronger model overall, scoring 34.0 to 30.6 on the Noometry Index.

Last verified . 31 shared benchmarks.

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

o1-mini OpenAI

34.0

Rank #235 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Llama-3.3-70B-Instruct scores higher in 3 categories and o1-mini in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where o1-mini leads 35.4 to 15.3.
  • The biggest single-benchmark swing is MATH Level 5: 41.6% for Llama-3.3-70B-Instruct and 89.2% for o1-mini.
  • Llama-3.3-70B-Instruct has downloadable open weights; the other is API-only.

Side by side

Llama-3.3-70B-Instruct and o1-mini specifications
Llama-3.3-70B-Instructo1-mini
ProviderMetaOpenAI
Noometry Index30.634.0
Released2024-12-062024-09-12
WeightsOpenProprietary
Context window128K—
Max output4K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.32—
Results tracked4339

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1-mini leads

Llama-3.3-70B-Instruct: 31.0 (#290), o1-mini: 35.5 (#224)

Coding benchmarks
BenchmarkLlama-3.3-70B-Instructo1-mini
WeirdML14.4%36.3%
LiveBench Coding36.6%48%
LMArena Coding12681362
Aider Polyglot—32.9%
SciCode26%—
BigCodeBench Instruct46.9%—
BigCodeBench Complete57.5%—
HumanEval+—89%
MBPP+—78.8%

Agentic & Tool Use Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 25.8 (#105), o1-mini: 24.6 (#118)

Agentic & Tool Use benchmarks
BenchmarkLlama-3.3-70B-Instructo1-mini
Berkeley Function Calling Leaderboard31.9%—
Cybench—10%
BALROG23%—

Reasoning Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 14.1 (#327), o1-mini: 8.8 (#346)

Reasoning benchmarks
BenchmarkLlama-3.3-70B-Instructo1-mini
SimpleBench19.9%18.1%
LiveBench Reasoning50.8%72.3%
LMArena Hard Prompts12571333
LiveBench Data Analysis49.5%57.9%
Epoch Capabilities Index127.33135.82
LiveBench50.2%57.8%
ARC-AGI-2—0.8%
ARC-AGI-1—14%
CritPt0%—
DTBench59.5%—
LMCA17.5%—
ForecastBench58.6—

Math o1-mini leads

Llama-3.3-70B-Instruct: 15.3 (#298), o1-mini: 35.4 (#186)

Math benchmarks
BenchmarkLlama-3.3-70B-Instructo1-mini
OTIS Mock AIME 2024-20255.1%46.9%
LiveBench Math42.2%62%
LMArena Math12671358
MATH Level 541.6%89.2%
FrontierMath (Feb 2025 set)—1.7%

Knowledge o1-mini leads

Llama-3.3-70B-Instruct: 30.6 (#226), o1-mini: 34.9 (#192)

Knowledge benchmarks
BenchmarkLlama-3.3-70B-Instructo1-mini
GPQA Diamond47.4%62.4%
Confabulations22.8%18.6%
LMArena Expert12251316
Vectara Hallucination Rate4.1%—
MMLU86.3%—

Multilingual o1-mini leads

Llama-3.3-70B-Instruct: 39.9 (#220), o1-mini: 43.6 (#182)

Multilingual benchmarks
BenchmarkLlama-3.3-70B-Instructo1-mini
LMArena Non-English12361289
LMArena Chinese12171314
LMArena French12811293
LMArena German12511278
LMArena Japanese11501245
LMArena Korean11431223
LMArena Russian12521283
LMArena Spanish12701303

Instruction Following Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 71.1 (#157), o1-mini: 66.7 (#206)

Instruction Following benchmarks
BenchmarkLlama-3.3-70B-Instructo1-mini
LiveBench Instruction Following82.7%65.4%
LMArena Instruction Following12421304

Long Context o1-mini leads

Llama-3.3-70B-Instruct: 26.4 (#295), o1-mini: 40.1 (#161)

Long Context benchmarks
BenchmarkLlama-3.3-70B-Instructo1-mini
LMArena Longer Query12561320
Fiction.LiveBench33.3%—

Writing & Preference Too close to call

Llama-3.3-70B-Instruct: 47.6 (#207), o1-mini: 48.4 (#202)

Writing & Preference benchmarks
BenchmarkLlama-3.3-70B-Instructo1-mini
LMArena Text12741317
LMArena Creative Writing12501244
LMArena Multi-Turn12801314
LiveBench Language39.2%40.9%
Short-Story Creative Writing—64.9%

Frequently asked questions

Is Llama-3.3-70B-Instruct better than o1-mini?

o1-mini is the stronger model overall, scoring 34.0 to 30.6 on the Noometry Index.

Is Llama-3.3-70B-Instruct or o1-mini better for coding?

o1-mini scores higher on coding benchmarks: 35.5 versus 31.0 in the Noometry coding category.

How many benchmarks do Llama-3.3-70B-Instruct and o1-mini share?

31 benchmarks have published results for both models. Llama-3.3-70B-Instruct has 43 scored results on Noometry and o1-mini has 39.

Related comparisons

Go deeper