Model comparison

o1-mini vs Qwen3.6 Max Preview

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 34.0 on the Noometry Index.

Last verified . 19 shared benchmarks.

o1-mini OpenAI

34.0

Rank #235 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 19 benchmarks with published results for both. o1-mini scores higher in 0 categories and Qwen3.6 Max Preview in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.6 Max Preview leads 41.7 to 8.8.
  • The biggest single-benchmark swing is SimpleBench: 18.1% for o1-mini and 63% for Qwen3.6 Max Preview.

Side by side

o1-mini and Qwen3.6 Max Preview specifications
o1-miniQwen3.6 Max Preview
ProviderOpenAIAlibaba (Qwen)
Noometry Index34.051.5
Released2024-09-122026-04-20
WeightsProprietaryProprietary
Context window—262K
Max output—66K
Input $ / M tokens—$1.30
Output $ / M tokens—$7.80
Results tracked3929

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

o1-mini: 35.5 (#224), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
Benchmarko1-miniQwen3.6 Max Preview
LMArena Coding13621471
SWE-bench Verified—76.7%
Aider Polyglot32.9%—
LMArena WebDev—1482
WeirdML36.3%—
LiveBench Coding48%—
HumanEval+89%—
MBPP+78.8%—

Agentic & Tool Use Not comparable

o1-mini: 24.6 (#118), Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
Benchmarko1-miniQwen3.6 Max Preview
Cybench10%—
Vending-Bench 2—4,254

Reasoning Qwen3.6 Max Preview leads

o1-mini: 8.8 (#346), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
Benchmarko1-miniQwen3.6 Max Preview
SimpleBench18.1%63%
LMArena Hard Prompts13331457
Epoch Capabilities Index135.82149.24
ARC-AGI-20.8%—
NYT Connections (extended)—74.1%
ARC-AGI-114%—
Chess Puzzles—20%
LiveBench Reasoning72.3%—
Mystery Game Puzzles—19%
DTBench—87.2%
LiveBench Data Analysis57.9%—
LMCA—42.5%
LiveBench57.8%—

Math Qwen3.6 Max Preview leads

o1-mini: 35.4 (#186), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
Benchmarko1-miniQwen3.6 Max Preview
OTIS Mock AIME 2024-202546.9%91.1%
LMArena Math13581465
FrontierMath (Feb 2025 set)1.7%23.1%
LiveBench Math62%—
MATH Level 589.2%—
FrontierMath Tier 4 (v1)—4.2%

Knowledge Qwen3.6 Max Preview leads

o1-mini: 34.9 (#192), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
Benchmarko1-miniQwen3.6 Max Preview
GPQA Diamond62.4%87.4%
LMArena Expert13161478
SimpleQA Verified—52%
Confabulations18.6%—

Multilingual Qwen3.6 Max Preview leads

o1-mini: 43.6 (#182), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
Benchmarko1-miniQwen3.6 Max Preview
LMArena Non-English12891437
LMArena Chinese13141487
LMArena French12931449
LMArena Russian12831445
LMArena Spanish13031454
LMArena German1278—
LMArena Japanese1245—
LMArena Korean1223—

Instruction Following Qwen3.6 Max Preview leads

o1-mini: 66.7 (#206), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
Benchmarko1-miniQwen3.6 Max Preview
LMArena Instruction Following13041438
LiveBench Instruction Following65.4%—

Long Context Qwen3.6 Max Preview leads

o1-mini: 40.1 (#161), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
Benchmarko1-miniQwen3.6 Max Preview
LMArena Longer Query13201457

Writing & Preference Qwen3.6 Max Preview leads

o1-mini: 48.4 (#202), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
Benchmarko1-miniQwen3.6 Max Preview
LMArena Text13171447
LMArena Creative Writing12441435
LMArena Multi-Turn13141456
Short-Story Creative Writing64.9%—
LiveBench Language40.9%—

Frequently asked questions

Is o1-mini better than Qwen3.6 Max Preview?

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 34.0 on the Noometry Index.

Is o1-mini or Qwen3.6 Max Preview better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 35.5 in the Noometry coding category.

How many benchmarks do o1-mini and Qwen3.6 Max Preview share?

19 benchmarks have published results for both models. o1-mini has 39 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper