Model comparison

o1-mini vs Qwen2.5 7B Instruct

o1-mini is the stronger model overall, scoring 34.0 to 29.0 on the Noometry Index.

Last verified . 3 shared benchmarks.

o1-mini OpenAI

34.0

Rank #235 Confirmed

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • They share 3 benchmarks with published results for both. o1-mini scores higher in 4 categories and Qwen2.5 7B Instruct in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where o1-mini leads 35.4 to 12.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 46.9% for o1-mini and 2.5% for Qwen2.5 7B Instruct.
  • Qwen2.5 7B Instruct has downloadable open weights; the other is API-only.

Side by side

o1-mini and Qwen2.5 7B Instruct specifications
o1-miniQwen2.5 7B Instruct
ProviderOpenAIAlibaba (Qwen)
Noometry Index34.029.0
Released2024-09-122024-09
WeightsProprietaryOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.17
Output $ / M tokens—$0.70
Results tracked3915

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 7B Instruct leads

o1-mini: 35.5 (#224), Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
Benchmarko1-miniQwen2.5 7B Instruct
Aider Polyglot32.9%—
WeirdML36.3%—
BigCodeBench Instruct—37.6%
LiveBench Coding48%—
LMArena Coding1362—
BigCodeBench Complete—46.1%
HumanEval+89%—
MBPP+78.8%—

Agentic & Tool Use Too close to call

o1-mini: 24.6 (#118), Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
Benchmarko1-miniQwen2.5 7B Instruct
Cybench10%—
BALROG—7.8%

Reasoning Qwen2.5 7B Instruct leads

o1-mini: 8.8 (#346), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
Benchmarko1-miniQwen2.5 7B Instruct
Epoch Capabilities Index135.82118.51
ARC-AGI-20.8%—
SimpleBench18.1%—
ARC-AGI-114%—
Chess Puzzles—0%
LiveBench Reasoning72.3%—
LMArena Hard Prompts1333—
DTBench—47.7%
LiveBench Data Analysis57.9%—
LMCA—6.4%
LiveBench57.8%—

Math o1-mini leads

o1-mini: 35.4 (#186), Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
Benchmarko1-miniQwen2.5 7B Instruct
OTIS Mock AIME 2024-202546.9%2.5%
Omni-MATH—29.4%
LiveBench Math62%—
LMArena Math1358—
MATH Level 589.2%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge o1-mini leads

o1-mini: 34.9 (#192), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
Benchmarko1-miniQwen2.5 7B Instruct
GPQA Diamond62.4%35.5%
MMLU-Pro—53.9%
Confabulations18.6%—
GPQA (HELM)—34.1%
LMArena Expert1316—
MMLU—72.9%

Multilingual Not comparable

o1-mini: 43.6 (#182), Qwen2.5 7B Instruct: —

Multilingual benchmarks
Benchmarko1-miniQwen2.5 7B Instruct
LMArena Non-English1289—
LMArena Chinese1314—
LMArena French1293—
LMArena German1278—
LMArena Japanese1245—
LMArena Korean1223—
LMArena Russian1283—
LMArena Spanish1303—

Instruction Following o1-mini leads

o1-mini: 66.7 (#206), Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
Benchmarko1-miniQwen2.5 7B Instruct
LiveBench Instruction Following65.4%—
IFEval—74.1%
LMArena Instruction Following1304—

Long Context Not comparable

o1-mini: 40.1 (#161), Qwen2.5 7B Instruct: —

Long Context benchmarks
Benchmarko1-miniQwen2.5 7B Instruct
LMArena Longer Query1320—

Writing & Preference Too close to call

o1-mini: 48.4 (#202), Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
Benchmarko1-miniQwen2.5 7B Instruct
LMArena Text1317—
LMArena Creative Writing1244—
Short-Story Creative Writing64.9%—
WildBench—73.1%
LMArena Multi-Turn1314—
LiveBench Language40.9%—

Frequently asked questions

Is o1-mini better than Qwen2.5 7B Instruct?

o1-mini is the stronger model overall, scoring 34.0 to 29.0 on the Noometry Index.

Is o1-mini or Qwen2.5 7B Instruct better for coding?

Qwen2.5 7B Instruct scores higher on coding benchmarks: 36.5 versus 35.5 in the Noometry coding category.

How many benchmarks do o1-mini and Qwen2.5 7B Instruct share?

3 benchmarks have published results for both models. o1-mini has 39 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper