Model comparison

o4-mini vs Qwen3.6 27B

o4-mini and Qwen3.6 27B score almost the same on the Noometry Index (41.6 vs 42.2), so choose on price, context window or the category you care about most.

Last verified . 9 shared benchmarks.

o4-mini OpenAI

41.6

Rank #132 Confirmed

Qwen3.6 27B Alibaba (Qwen)

42.2

Rank #117 Confirmed

Summary

  • They share 9 benchmarks with published results for both. o4-mini scores higher in 2 categories and Qwen3.6 27B in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.6 27B leads 52.4 to 43.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 81.7% for o4-mini and 91.1% for Qwen3.6 27B.
  • Qwen3.6 27B is cheaper at $0.60 / $3.60 per million input/output tokens, against $1.10 / $4.40 for o4-mini.
  • Qwen3.6 27B accepts more context: 262K tokens versus 200K.
  • Qwen3.6 27B has downloadable open weights; the other is API-only.

Side by side

o4-mini and Qwen3.6 27B specifications
o4-miniQwen3.6 27B
ProviderOpenAIAlibaba (Qwen)
Noometry Index41.642.2
Released2025-04-162026-04-22
WeightsProprietaryOpen
Context window200K262K
Max output100K66K
Input $ / M tokens$1.10$0.60
Output $ / M tokens$4.40$3.60
Results tracked6011

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o4-mini leads

o4-mini: 40.9 (#127), Qwen3.6 27B: 39.1 (#163)

Coding benchmarks
Benchmarko4-miniQwen3.6 27B
SWE-bench Verified (bash only)45%—
Aider Polyglot72%—
SciCode—37.3%
GSO3.6%—
WeirdML52.6%—
LMArena Coding1368—
CadEval62%—
ALE-Bench826.17—
AlgoTune1.72—

Agentic & Tool Use Not comparable

o4-mini: 32.6 (#61), Qwen3.6 27B: —

Agentic & Tool Use benchmarks
Benchmarko4-miniQwen3.6 27B
Berkeley Function Calling Leaderboard53.2%—
GDPval25.3%—
METR Time Horizons63.9%—

Reasoning Too close to call

o4-mini: 24.6 (#162), Qwen3.6 27B: 25.0 (#153)

Reasoning benchmarks
Benchmarko4-miniQwen3.6 27B
CritPt0.6%0.9%
Chess Puzzles26%22%
Mystery Game Puzzles5%7%
DTBench77.6%78.1%
LMCA26.5%34.5%
Epoch Capabilities Index145.64146.5
ARC-AGI-26.1%—
SimpleBench38.7%—
Kagi LLM Benchmark67.6%—
ARC-AGI-158.7%—
EnigmaEval9.2%—
LMArena Hard Prompts1351—
ForecastBench61.8—

Math Qwen3.6 27B leads

o4-mini: 40.8 (#89), Qwen3.6 27B: 48.5 (#62)

Math benchmarks
Benchmarko4-miniQwen3.6 27B
FrontierMath (Tiers 1-3)36.1%35.1%
OTIS Mock AIME 2024-202581.7%91.1%
FrontierMath Tier 44.9%—
Omni-MATH72%—
LMArena Math1389—
MATH Level 597.8%—
FrontierMath (Feb 2025 set)24.8%—
FrontierMath Tier 4 (v1)6.3%—

Knowledge Qwen3.6 27B leads

o4-mini: 43.6 (#91), Qwen3.6 27B: 52.4 (#63)

Knowledge benchmarks
Benchmarko4-miniQwen3.6 27B
GPQA Diamond79.6%85.9%
Humanity's Last Exam18.1%—
SimpleQA Verified19.6%—
MMLU-Pro82%—
Confabulations15.8%—
Vectara Hallucination Rate18.6%—
GPQA (HELM)73.5%—
LMArena Expert1343—

Multimodal Not comparable

o4-mini: 40.2 (#49), Qwen3.6 27B: —

Multimodal benchmarks
Benchmarko4-miniQwen3.6 27B
LMArena Vision1194—
GeoBench64%—
VPCT57.5%—

Multilingual Not comparable

o4-mini: 47.0 (#154), Qwen3.6 27B: —

Multilingual benchmarks
Benchmarko4-miniQwen3.6 27B
LMArena Non-English1337—
LMArena Chinese1354—
LMArena French1364—
LMArena German1336—
LMArena Japanese1308—
LMArena Korean1312—
LMArena Russian1334—
LMArena Spanish1347—

Instruction Following Not comparable

o4-mini: 75.2 (#68), Qwen3.6 27B: —

Instruction Following benchmarks
Benchmarko4-miniQwen3.6 27B
IFEval92.8%—
LMArena Instruction Following1321—

Long Context Not comparable

o4-mini: 45.5 (#33), Qwen3.6 27B: —

Long Context benchmarks
Benchmarko4-miniQwen3.6 27B
Fiction.LiveBench77.8%—
LMArena Longer Query1315—

Writing & Preference o4-mini leads

o4-mini: 54.0 (#152), Qwen3.6 27B: 50.3 (#181)

Writing & Preference benchmarks
Benchmarko4-miniQwen3.6 27B
LMArena Text1353—
LMArena Creative Writing1294—
Short-Story Creative Writing75%—
WildBench85.4%—
EQ-Bench 4—1026
LMArena Multi-Turn1350—

Frequently asked questions

Is o4-mini better than Qwen3.6 27B?

o4-mini and Qwen3.6 27B score almost the same on the Noometry Index (41.6 vs 42.2), so choose on price, context window or the category you care about most.

Which is cheaper, o4-mini or Qwen3.6 27B?

Qwen3.6 27B is cheaper. It lists at $0.60 per million input tokens and $3.60 per million output tokens; o4-mini lists at $1.10 and $4.40.

Is o4-mini or Qwen3.6 27B better for coding?

o4-mini scores higher on coding benchmarks: 40.9 versus 39.1 in the Noometry coding category.

Which has the bigger context window?

Qwen3.6 27B does, with 262K tokens against 200K.

How many benchmarks do o4-mini and Qwen3.6 27B share?

9 benchmarks have published results for both models. o4-mini has 60 scored results on Noometry and Qwen3.6 27B has 11.

Related comparisons

Go deeper