Model comparison

GPT-5 Mini vs Qwen3.5-Flash

GPT-5 Mini and Qwen3.5-Flash score almost the same on the Noometry Index (41.8 vs 42.5), so choose on price, context window or the category you care about most.

Last verified . 31 shared benchmarks.

GPT-5 Mini OpenAI

41.8

Rank #128 Confirmed

Qwen3.5-Flash Alibaba (Qwen)

42.5

Rank #112 Confirmed

Summary

  • They share 31 benchmarks with published results for both. GPT-5 Mini scores higher in 4 categories and Qwen3.5-Flash in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.5-Flash leads 33.7 to 23.9.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 46.7% for GPT-5 Mini and 18.2% for Qwen3.5-Flash.
  • Qwen3.5-Flash is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.25 / $2 for GPT-5 Mini.
  • Qwen3.5-Flash accepts more context: 1M tokens versus 400K.

Side by side

GPT-5 Mini and Qwen3.5-Flash specifications
GPT-5 MiniQwen3.5-Flash
ProviderOpenAIAlibaba (Qwen)
Noometry Index41.842.5
Released2025-08-072026-02-23
WeightsProprietaryProprietary
Context window400K1M
Max output128K66K
Input $ / M tokens$0.25$0.10
Output $ / M tokens$2$0.40
Results tracked6032

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5 Mini leads

GPT-5 Mini: 40.1 (#146), Qwen3.5-Flash: 34.2 (#242)

Coding benchmarks
BenchmarkGPT-5 MiniQwen3.5-Flash
LMArena Coding14061412
ALE-Bench799.77221.8
SWE-bench Verified64.7%—
SWE-bench Verified (bash only)59.8%—
LMArena WebDev—1244
SWE-bench Multilingual39.7%—
SciCode39.2%—
WeirdML52.7%—
AlgoTune1.38—

Agentic & Tool Use Not comparable

GPT-5 Mini: 31.1 (#70), Qwen3.5-Flash: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5 MiniQwen3.5-Flash
Vending-Bench 2-31.18462.69
Terminal-Bench34.8%—
Berkeley Function Calling Leaderboard55.5%—

Reasoning Qwen3.5-Flash leads

GPT-5 Mini: 23.9 (#168), Qwen3.5-Flash: 33.7 (#72)

Reasoning benchmarks
BenchmarkGPT-5 MiniQwen3.5-Flash
Chess Puzzles30%21%
LMArena Hard Prompts13801403
Mystery Game Puzzles10%20%
DTBench80.5%82.9%
LMCA34.2%29.1%
Epoch Capabilities Index145.52143.98
ARC-AGI-24.4%—
Kagi LLM Benchmark70.3%—
ARC-AGI-154.3%—
CritPt0%—
EnigmaEval8.2%—
ForecastBench61—

Math GPT-5 Mini leads

GPT-5 Mini: 46.7 (#69), Qwen3.5-Flash: 37.4 (#158)

Math benchmarks
BenchmarkGPT-5 MiniQwen3.5-Flash
FrontierMath (Tiers 1-3)46.7%18.2%
OTIS Mock AIME 2024-202586.7%84.4%
LMArena Math13781407
FrontierMath (Feb 2025 set)27.2%6.2%
FrontierMath Tier 4 (v1)6.3%0%
FrontierMath Tier 412.2%—
ProofBench9%—
Omni-MATH72.2%—
MATH Level 597.8%—

Knowledge GPT-5 Mini leads

GPT-5 Mini: 45.6 (#86), Qwen3.5-Flash: 43.2 (#93)

Knowledge benchmarks
BenchmarkGPT-5 MiniQwen3.5-Flash
GPQA Diamond75%82.3%
SimpleQA Verified21.6%20.3%
Vectara Hallucination Rate12.9%10.5%
LMArena Expert13791407
Humanity's Last Exam19.4%—
MMLU-Pro83.5%—
Confabulations13.3%—
GPQA (HELM)75.6%—

Multimodal Not comparable

GPT-5 Mini: 35.6 (#85), Qwen3.5-Flash: —

Multimodal benchmarks
BenchmarkGPT-5 MiniQwen3.5-Flash
LMArena Vision1202—
VPCT40.2%—

Multilingual Qwen3.5-Flash leads

GPT-5 Mini: 48.9 (#137), Qwen3.5-Flash: 50.5 (#121)

Multilingual benchmarks
BenchmarkGPT-5 MiniQwen3.5-Flash
LMArena Non-English13631385
LMArena Chinese13851446
LMArena French13861412
LMArena German13661390
LMArena Japanese13411368
LMArena Korean13081344
LMArena Russian13621379
LMArena Spanish13551400

Instruction Following GPT-5 Mini leads

GPT-5 Mini: 76.2 (#46), Qwen3.5-Flash: 72.6 (#139)

Instruction Following benchmarks
BenchmarkGPT-5 MiniQwen3.5-Flash
LMArena Instruction Following13571374
IFEval92.7%—

Long Context Too close to call

GPT-5 Mini: 41.9 (#132), Qwen3.5-Flash: 42.4 (#124)

Long Context benchmarks
BenchmarkGPT-5 MiniQwen3.5-Flash
LMArena Longer Query13551392
Fiction.LiveBench69.4%—

Writing & Preference Qwen3.5-Flash leads

GPT-5 Mini: 55.2 (#148), Qwen3.5-Flash: 57.9 (#122)

Writing & Preference benchmarks
BenchmarkGPT-5 MiniQwen3.5-Flash
LMArena Text13731397
LMArena Creative Writing13251343
LMArena Multi-Turn13631393
Short-Story Creative Writing83.1%—
EQ-Bench Creative Writing1313—
WildBench85.5%—

Frequently asked questions

Is GPT-5 Mini better than Qwen3.5-Flash?

GPT-5 Mini and Qwen3.5-Flash score almost the same on the Noometry Index (41.8 vs 42.5), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5 Mini or Qwen3.5-Flash?

Qwen3.5-Flash is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; GPT-5 Mini lists at $0.25 and $2.

Is GPT-5 Mini or Qwen3.5-Flash better for coding?

GPT-5 Mini scores higher on coding benchmarks: 40.1 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

Qwen3.5-Flash does, with 1M tokens against 400K.

How many benchmarks do GPT-5 Mini and Qwen3.5-Flash share?

31 benchmarks have published results for both models. GPT-5 Mini has 60 scored results on Noometry and Qwen3.5-Flash has 32.

Related comparisons

Go deeper