Model comparison

o4-mini vs Qwen3.5 27B

o4-mini and Qwen3.5 27B score almost the same on the Noometry Index (41.6 vs 41.9), so choose on price, context window or the category you care about most.

Last verified . 23 shared benchmarks.

o4-mini OpenAI

41.6

Rank #132 Confirmed

Qwen3.5 27B Alibaba (Qwen)

41.9

Rank #127 Confirmed

Summary

  • They share 23 benchmarks with published results for both. o4-mini scores higher in 6 categories and Qwen3.5 27B in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o4-mini leads 43.6 to 38.0.
  • The biggest single-benchmark swing is WeirdML: 52.6% for o4-mini and 39.5% for Qwen3.5 27B.
  • Qwen3.5 27B is cheaper at $0.30 / $2.40 per million input/output tokens, against $1.10 / $4.40 for o4-mini.
  • Qwen3.5 27B accepts more context: 262K tokens versus 200K.
  • Qwen3.5 27B has downloadable open weights; the other is API-only.

Side by side

o4-mini and Qwen3.5 27B specifications
o4-miniQwen3.5 27B
ProviderOpenAIAlibaba (Qwen)
Noometry Index41.641.9
Released2025-04-162026-02-23
WeightsProprietaryOpen
Context window200K262K
Max output100K66K
Input $ / M tokens$1.10$0.30
Output $ / M tokens$4.40$2.40
Results tracked6028

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o4-mini leads

o4-mini: 40.9 (#127), Qwen3.5 27B: 38.9 (#168)

Coding benchmarks
Benchmarko4-miniQwen3.5 27B
WeirdML52.6%39.5%
LMArena Coding13681427
ALE-Bench826.17349.45
SWE-bench Verified (bash only)45%—
Aider Polyglot72%—
LMArena WebDev—1358
GSO3.6%—
CadEval62%—
AlgoTune1.72—

Agentic & Tool Use Not comparable

o4-mini: 32.6 (#61), Qwen3.5 27B: —

Agentic & Tool Use benchmarks
Benchmarko4-miniQwen3.5 27B
Berkeley Function Calling Leaderboard53.2%—
GDPval25.3%—
METR Time Horizons63.9%—
Vending-Bench 2—201.98

Reasoning Qwen3.5 27B leads

o4-mini: 24.6 (#162), Qwen3.5 27B: 27.5 (#117)

Reasoning benchmarks
Benchmarko4-miniQwen3.5 27B
LMArena Hard Prompts13511414
DTBench77.6%82.4%
LMCA26.5%34%
ARC-AGI-26.1%—
SimpleBench38.7%—
Kagi LLM Benchmark67.6%—
NYT Connections (extended)—47.9%
ARC-AGI-158.7%—
CritPt0.6%—
Chess Puzzles26%—
EnigmaEval9.2%—
Thematic Generalization—45.5%
Mystery Game Puzzles5%—
Epoch Capabilities Index145.64—
ForecastBench61.8—

Math o4-mini leads

o4-mini: 40.8 (#89), Qwen3.5 27B: 38.8 (#127)

Knowledge o4-mini leads

o4-mini: 43.6 (#91), Qwen3.5 27B: 38.0 (#150)

Knowledge benchmarks
Benchmarko4-miniQwen3.5 27B
Vectara Hallucination Rate18.6%12.1%
LMArena Expert13431428
GPQA Diamond79.6%—
Humanity's Last Exam18.1%—
SimpleQA Verified19.6%—
MMLU-Pro82%—
Confabulations15.8%—
GPQA (HELM)73.5%—

Multimodal Too close to call

o4-mini: 40.2 (#49), Qwen3.5 27B: 39.4 (#59)

Multimodal benchmarks
Benchmarko4-miniQwen3.5 27B
LMArena Vision11941241
GeoBench64%—
VPCT57.5%—

Multilingual Qwen3.5 27B leads

o4-mini: 47.0 (#154), Qwen3.5 27B: 50.8 (#115)

Multilingual benchmarks
Benchmarko4-miniQwen3.5 27B
LMArena Non-English13371390
LMArena Chinese13541478
LMArena French13641410
LMArena German13361393
LMArena Japanese13081345
LMArena Korean13121358
LMArena Russian13341390
LMArena Spanish13471407

Instruction Following o4-mini leads

o4-mini: 75.2 (#68), Qwen3.5 27B: 73.5 (#119)

Instruction Following benchmarks
Benchmarko4-miniQwen3.5 27B
LMArena Instruction Following13211393
IFEval92.8%—

Long Context o4-mini leads

o4-mini: 45.5 (#33), Qwen3.5 27B: 43.1 (#106)

Long Context benchmarks
Benchmarko4-miniQwen3.5 27B
LMArena Longer Query13151413
Fiction.LiveBench77.8%—

Writing & Preference Qwen3.5 27B leads

o4-mini: 54.0 (#152), Qwen3.5 27B: 59.3 (#111)

Writing & Preference benchmarks
Benchmarko4-miniQwen3.5 27B
LMArena Text13531409
LMArena Creative Writing12941362
LMArena Multi-Turn13501410
Short-Story Creative Writing75%—
WildBench85.4%—

Frequently asked questions

Is o4-mini better than Qwen3.5 27B?

o4-mini and Qwen3.5 27B score almost the same on the Noometry Index (41.6 vs 41.9), so choose on price, context window or the category you care about most.

Which is cheaper, o4-mini or Qwen3.5 27B?

Qwen3.5 27B is cheaper. It lists at $0.30 per million input tokens and $2.40 per million output tokens; o4-mini lists at $1.10 and $4.40.

Is o4-mini or Qwen3.5 27B better for coding?

o4-mini scores higher on coding benchmarks: 40.9 versus 38.9 in the Noometry coding category.

Which has the bigger context window?

Qwen3.5 27B does, with 262K tokens against 200K.

How many benchmarks do o4-mini and Qwen3.5 27B share?

23 benchmarks have published results for both models. o4-mini has 60 scored results on Noometry and Qwen3.5 27B has 28.

Related comparisons

Go deeper