Model comparison
o3-pro vs Qwen3.6 27B
o3-pro and Qwen3.6 27B score almost the same on the Noometry Index (42.9 vs 42.2), so choose on price, context window or the category you care about most.
Last verified . 3 shared benchmarks.
Summary
- They share 3 benchmarks with published results for both. o3-pro scores higher in 2 categories and Qwen3.6 27B in 2 categories; 4 gaps are clear of the uncertainty.
- The widest gap is in knowledge, where Qwen3.6 27B leads 52.4 to 29.5.
- The biggest single-benchmark swing is DTBench: 86.9% for o3-pro and 78.1% for Qwen3.6 27B.
- Qwen3.6 27B is cheaper at $0.60 / $3.60 per million input/output tokens, against $20 / $80 for o3-pro.
- Qwen3.6 27B accepts more context: 262K tokens versus 200K.
- Qwen3.6 27B has downloadable open weights; the other is API-only.
Side by side
| o3-pro | Qwen3.6 27B | |
|---|---|---|
| Provider | OpenAI | Alibaba (Qwen) |
| Noometry Index | 42.9 | 42.2 |
| Released | 2025-06-10 | 2026-04-22 |
| Weights | Proprietary | Open |
| Context window | 200K | 262K |
| Max output | 100K | 66K |
| Input $ / M tokens | $20 | $0.60 |
| Output $ / M tokens | $80 | $3.60 |
| Results tracked | 12 | 11 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding o3-pro leads
o3-pro: 55.5 (#24), Qwen3.6 27B: 39.1 (#163)
| Benchmark | o3-pro | Qwen3.6 27B |
|---|---|---|
| Aider Polyglot | 84.9% | — |
| SciCode | — | 37.3% |
| WeirdML | 58.2% | — |
Reasoning Qwen3.6 27B leads
o3-pro: 23.8 (#171), Qwen3.6 27B: 25.0 (#153)
| Benchmark | o3-pro | Qwen3.6 27B |
|---|---|---|
| DTBench | 86.9% | 78.1% |
| LMCA | 38.5% | 34.5% |
| Epoch Capabilities Index | 147.42 | 146.5 |
| ARC-AGI-2 | 4.9% | — |
| Kagi LLM Benchmark | 72.1% | — |
| ARC-AGI-1 | 59.3% | — |
| CritPt | — | 0.9% |
| Chess Puzzles | — | 22% |
| Mystery Game Puzzles | — | 7% |
Math Not comparable
o3-pro: —, Qwen3.6 27B: 48.5 (#62)
| Benchmark | o3-pro | Qwen3.6 27B |
|---|---|---|
| FrontierMath (Tiers 1-3) | — | 35.1% |
| OTIS Mock AIME 2024-2025 | — | 91.1% |
Knowledge Qwen3.6 27B leads
o3-pro: 29.5 (#238), Qwen3.6 27B: 52.4 (#63)
| Benchmark | o3-pro | Qwen3.6 27B |
|---|---|---|
| GPQA Diamond | — | 85.9% |
| Confabulations | 14.2% | — |
| Vectara Hallucination Rate | 23.3% | — |
Long Context Not comparable
o3-pro: 72.2 (#1), Qwen3.6 27B: —
| Benchmark | o3-pro | Qwen3.6 27B |
|---|---|---|
| Fiction.LiveBench | 97.2% | — |
Writing & Preference o3-pro leads
o3-pro: 57.1 (#133), Qwen3.6 27B: 50.3 (#181)
| Benchmark | o3-pro | Qwen3.6 27B |
|---|---|---|
| Short-Story Creative Writing | 84.4% | — |
| EQ-Bench 4 | — | 1026 |
Frequently asked questions
Is o3-pro better than Qwen3.6 27B?
o3-pro and Qwen3.6 27B score almost the same on the Noometry Index (42.9 vs 42.2), so choose on price, context window or the category you care about most.
Which is cheaper, o3-pro or Qwen3.6 27B?
Qwen3.6 27B is cheaper. It lists at $0.60 per million input tokens and $3.60 per million output tokens; o3-pro lists at $20 and $80.
Is o3-pro or Qwen3.6 27B better for coding?
o3-pro scores higher on coding benchmarks: 55.5 versus 39.1 in the Noometry coding category.
Which has the bigger context window?
Qwen3.6 27B does, with 262K tokens against 200K.
How many benchmarks do o3-pro and Qwen3.6 27B share?
3 benchmarks have published results for both models. o3-pro has 12 scored results on Noometry and Qwen3.6 27B has 11.