Model comparison

o3 vs Qwen Turbo

o3 is the stronger model overall, scoring 47.5 to 27.1 on the Noometry Index. Qwen Turbo costs 40× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Last verified . 3 shared benchmarks.

o3 OpenAI

47.5

Rank #61 Confirmed

Qwen Turbo Alibaba (Qwen)

27.1

Rank #335 Reported

Summary

  • They share 3 benchmarks with published results for both. o3 scores higher in 2 categories and Qwen Turbo in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in math, where o3 leads 50.2 to 15.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 84.4% for o3 and 6.1% for Qwen Turbo.
  • Qwen Turbo is cheaper at $0.05 / $0.20 per million input/output tokens, against $2 / $8 for o3.
  • Qwen Turbo accepts more context: 1M tokens versus 200K.

Side by side

o3 and Qwen Turbo specifications
o3Qwen Turbo
ProviderOpenAIAlibaba (Qwen)
Noometry Index47.527.1
Released2025-04-162024-11-01
WeightsProprietaryProprietary
Context window200K1M
Max output100K16K
Input $ / M tokens$2$0.05
Output $ / M tokens$8$0.20
Results tracked633

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

o3: 46.8 (#64), Qwen Turbo: —

Coding benchmarks
Benchmarko3Qwen Turbo
SWE-bench Verified62.3%—
SWE-bench Verified (bash only)58.4%—
Aider Polyglot81.3%—
GSO8.8%—
WeirdML52.4%—
LMArena Coding1408—
CadEval74%—
ALE-Bench933.55—

Agentic & Tool Use Not comparable

o3: 34.5 (#44), Qwen Turbo: —

Agentic & Tool Use benchmarks
Benchmarko3Qwen Turbo
Berkeley Function Calling Leaderboard63%—
GDPval30.8%—
DeepResearch Bench45.2%—
OSWorld23%—
LMArena Search1144—
METR Time Horizons65.4%—

Reasoning Not comparable

o3: 32.0 (#78), Qwen Turbo: —

Reasoning benchmarks
Benchmarko3Qwen Turbo
ARC-AGI-26.5%—
SimpleBench53.1%—
Kagi LLM Benchmark67.6%—
ARC-AGI-160.8%—
CritPt1.4%—
Chess Puzzles38%—
EnigmaEval13.1%—
LMArena Hard Prompts1402—
Mystery Game Puzzles29%—
DTBench84.8%—
LMCA39.7%—
Epoch Capabilities Index146.86—
ForecastBench62.5—

Math o3 leads

o3: 50.2 (#58), Qwen Turbo: 15.3 (#297)

Math benchmarks
Benchmarko3Qwen Turbo
OTIS Mock AIME 2024-202584.4%6.1%
MATH Level 597.8%56.2%
FrontierMath (Tiers 1-3)33.3%—
Omni-MATH71.4%—
LMArena Math1426—
FrontierMath (Feb 2025 set)18.7%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge o3 leads

o3: 54.6 (#52), Qwen Turbo: 22.2 (#272)

Knowledge benchmarks
Benchmarko3Qwen Turbo
GPQA Diamond81.8%41.8%
Humanity's Last Exam20.3%—
SimpleQA Verified49.4%—
MMLU-Pro85.9%—
Confabulations14.4%—
GPQA (HELM)75.3%—
LMArena Expert1402—

Multimodal Not comparable

o3: 41.4 (#36), Qwen Turbo: —

Multimodal benchmarks
Benchmarko3Qwen Turbo
LMArena Vision1214—
GeoBench74%—
VPCT52%—

Multilingual Not comparable

o3: 51.7 (#105), Qwen Turbo: —

Multilingual benchmarks
Benchmarko3Qwen Turbo
LMArena Non-English1401—
LMArena Chinese1437—
LMArena French1430—
LMArena German1420—
LMArena Japanese1403—
LMArena Korean1370—
LMArena Russian1406—
LMArena Spanish1395—

Instruction Following Not comparable

o3: 72.8 (#127), Qwen Turbo: —

Instruction Following benchmarks
Benchmarko3Qwen Turbo
IFEval86.9%—
LMArena Instruction Following1368—

Long Context Not comparable

o3: 53.3 (#6), Qwen Turbo: —

Long Context benchmarks
Benchmarko3Qwen Turbo
Fiction.LiveBench88.9%—
CL-bench17.8%—
LMArena Longer Query1372—

Writing & Preference Not comparable

o3: 63.5 (#64), Qwen Turbo: —

Writing & Preference benchmarks
Benchmarko3Qwen Turbo
LMArena Text1410—
LMArena Creative Writing1359—
Short-Story Creative Writing83.9%—
EQ-Bench Creative Writing1676—
WildBench86.1%—
LMArena Multi-Turn1405—

Frequently asked questions

Is o3 better than Qwen Turbo?

o3 is the stronger model overall, scoring 47.5 to 27.1 on the Noometry Index. Qwen Turbo costs 40× less per token, which makes it the better buy when o3's lead doesn't matter for your workload.

Which is cheaper, o3 or Qwen Turbo?

Qwen Turbo is cheaper. It lists at $0.05 per million input tokens and $0.20 per million output tokens; o3 lists at $2 and $8.

Which has the bigger context window?

Qwen Turbo does, with 1M tokens against 200K.

How many benchmarks do o3 and Qwen Turbo share?

3 benchmarks have published results for both models. o3 has 63 scored results on Noometry and Qwen Turbo has 3.

Related comparisons

Go deeper