Model comparison

o3 vs Qwen2.5-Coder-32B

o3 is the stronger model overall, scoring 47.5 to 33.4 on the Noometry Index. Qwen2.5-Coder-32B costs 4.7× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Last verified . 15 shared benchmarks.

o3 OpenAI

47.5

Rank #61 Confirmed

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Summary

  • They share 15 benchmarks with published results for both. o3 scores higher in 8 categories and Qwen2.5-Coder-32B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where o3 leads 46.8 to 22.6.
  • The biggest single-benchmark swing is Aider Polyglot: 81.3% for o3 and 16.4% for Qwen2.5-Coder-32B.
  • Qwen2.5-Coder-32B is cheaper at $0.66 / $1 per million input/output tokens, against $2 / $8 for o3.
  • o3 accepts more context: 200K tokens versus 33K.
  • Qwen2.5-Coder-32B has downloadable open weights; the other is API-only.

Side by side

o3 and Qwen2.5-Coder-32B specifications
o3Qwen2.5-Coder-32B
ProviderOpenAIAlibaba (Qwen)
Noometry Index47.533.4
Released2025-04-162024-09-18
WeightsProprietaryOpen
Context window200K33K
Max output100K29K
Input $ / M tokens$2$0.66
Output $ / M tokens$8$1
Results tracked6331

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

o3: 46.8 (#64), Qwen2.5-Coder-32B: 22.6 (#333)

Coding benchmarks
Benchmarko3Qwen2.5-Coder-32B
SWE-bench Verified (bash only)58.4%9%
Aider Polyglot81.3%16.4%
LMArena Coding14081276
SWE-bench Verified62.3%—
GSO8.8%—
WeirdML52.4%—
BigCodeBench Instruct—49%
LiveBench Coding—56.9%
BigCodeBench Complete—58%
CadEval74%—
ALE-Bench933.55—
HumanEval+—87.2%
MBPP+—77%

Agentic & Tool Use Not comparable

o3: 34.5 (#44), Qwen2.5-Coder-32B: —

Agentic & Tool Use benchmarks
Benchmarko3Qwen2.5-Coder-32B
Berkeley Function Calling Leaderboard63%—
GDPval30.8%—
DeepResearch Bench45.2%—
OSWorld23%—
LMArena Search1144—
METR Time Horizons65.4%—

Reasoning o3 leads

o3: 32.0 (#78), Qwen2.5-Coder-32B: 21.2 (#225)

Reasoning benchmarks
Benchmarko3Qwen2.5-Coder-32B
LMArena Hard Prompts14021251
Epoch Capabilities Index146.86119.49
ARC-AGI-26.5%—
SimpleBench53.1%—
Kagi LLM Benchmark67.6%—
ARC-AGI-160.8%—
CritPt1.4%—
Chess Puzzles38%—
EnigmaEval13.1%—
LiveBench Reasoning—42.1%
Mystery Game Puzzles29%—
DTBench84.8%—
LiveBench Data Analysis—49.9%
LMCA39.7%—
ForecastBench62.5—
HellaSwag—83%
LiveBench—46.2%
WinoGrande—80.8%

Math o3 leads

o3: 50.2 (#58), Qwen2.5-Coder-32B: 33.3 (#204)

Math benchmarks
Benchmarko3Qwen2.5-Coder-32B
LMArena Math14261251
FrontierMath (Tiers 1-3)33.3%—
OTIS Mock AIME 2024-202584.4%—
Omni-MATH71.4%—
LiveBench Math—46.6%
MATH Level 597.8%—
FrontierMath (Feb 2025 set)18.7%—
FrontierMath Tier 4 (v1)2.1%—
GSM8K—93%

Knowledge o3 leads

o3: 54.6 (#52), Qwen2.5-Coder-32B: 33.4 (#203)

Knowledge benchmarks
Benchmarko3Qwen2.5-Coder-32B
LMArena Expert14021221
GPQA Diamond81.8%—
Humanity's Last Exam20.3%—
SimpleQA Verified49.4%—
MMLU-Pro85.9%—
Confabulations14.4%—
GPQA (HELM)75.3%—
ARC (AI2) Challenge—70.5%
MMLU—79.1%

Multimodal Not comparable

o3: 41.4 (#36), Qwen2.5-Coder-32B: —

Multimodal benchmarks
Benchmarko3Qwen2.5-Coder-32B
LMArena Vision1214—
GeoBench74%—
VPCT52%—

Multilingual o3 leads

o3: 51.7 (#105), Qwen2.5-Coder-32B: 37.8 (#235)

Multilingual benchmarks
Benchmarko3Qwen2.5-Coder-32B
LMArena Non-English14011205
LMArena Chinese14371222
LMArena Russian14061228
LMArena French1430—
LMArena German1420—
LMArena Japanese1403—
LMArena Korean1370—
LMArena Spanish1395—

Instruction Following o3 leads

o3: 72.8 (#127), Qwen2.5-Coder-32B: 61.4 (#245)

Instruction Following benchmarks
Benchmarko3Qwen2.5-Coder-32B
LMArena Instruction Following13681223
LiveBench Instruction Following—58.7%
IFEval86.9%—

Long Context o3 leads

o3: 53.3 (#6), Qwen2.5-Coder-32B: 38.0 (#208)

Long Context benchmarks
Benchmarko3Qwen2.5-Coder-32B
LMArena Longer Query13721251
Fiction.LiveBench88.9%—
CL-bench17.8%—

Writing & Preference o3 leads

o3: 63.5 (#64), Qwen2.5-Coder-32B: 41.6 (#240)

Writing & Preference benchmarks
Benchmarko3Qwen2.5-Coder-32B
LMArena Text14101230
LMArena Creative Writing13591174
LMArena Multi-Turn14051222
Short-Story Creative Writing83.9%—
EQ-Bench Creative Writing1676—
WildBench86.1%—
LiveBench Language—23.3%

Frequently asked questions

Is o3 better than Qwen2.5-Coder-32B?

o3 is the stronger model overall, scoring 47.5 to 33.4 on the Noometry Index. Qwen2.5-Coder-32B costs 4.7× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Which is cheaper, o3 or Qwen2.5-Coder-32B?

Qwen2.5-Coder-32B is cheaper. It lists at $0.66 per million input tokens and $1 per million output tokens; o3 lists at $2 and $8.

Is o3 or Qwen2.5-Coder-32B better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 22.6 in the Noometry coding category.

Which has the bigger context window?

o3 does, with 200K tokens against 33K.

How many benchmarks do o3 and Qwen2.5-Coder-32B share?

15 benchmarks have published results for both models. o3 has 63 scored results on Noometry and Qwen2.5-Coder-32B has 31.

Related comparisons

Go deeper