Model comparison

GPT-4o mini vs QwQ-32B

QwQ-32B is the stronger model overall, scoring 39.8 to 25.5 on the Noometry Index.

Last verified . 34 shared benchmarks.

GPT-4o mini OpenAI

25.5

Rank #343 Confirmed

QwQ-32B Alibaba (Qwen)

39.8

Rank #159 Confirmed

Summary

  • They share 34 benchmarks with published results for both. GPT-4o mini scores higher in 0 categories and QwQ-32B in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where QwQ-32B leads 38.0 to 10.4.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 6.9% for GPT-4o mini and 59.2% for QwQ-32B.
  • QwQ-32B has downloadable open weights; the other is API-only.

Side by side

GPT-4o mini and QwQ-32B specifications
GPT-4o miniQwQ-32B
ProviderOpenAIAlibaba (Qwen)
Noometry Index25.539.8
Released2024-07-182024-11-28
WeightsProprietaryOpen
Context window128K—
Max output16K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked6036

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding QwQ-32B leads

GPT-4o mini: 22.0 (#335), QwQ-32B: 35.4 (#226)

Coding benchmarks
BenchmarkGPT-4o miniQwQ-32B
Aider Polyglot3.6%20.9%
BigCodeBench Instruct46.1%44.6%
LiveBench Coding43.1%72.2%
LMArena Coding12901333
BigCodeBench Complete57.4%54.4%
WeirdML11.8%—
HumanEval+83.5%—
MBPP+72.2%—

Agentic & Tool Use Not comparable

GPT-4o mini: 27.5 (#101), QwQ-32B: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4o miniQwQ-32B
BALROG17.4%—

Reasoning QwQ-32B leads

GPT-4o mini: 8.7 (#347), QwQ-32B: 23.7 (#174)

Reasoning benchmarks
BenchmarkGPT-4o miniQwQ-32B
Chess Puzzles0%5%
LiveBench Reasoning32.8%83.5%
LMArena Hard Prompts12671325
LiveBench Data Analysis50%65%
Epoch Capabilities Index126.56137.6
LiveBench41.3%72%
ARC-AGI-20%—
SimpleBench10.7%—
Kagi LLM Benchmark28.8%—
Mystery Game Puzzles12%—
DTBench54.4%—
LMCA10.4%—
ForecastBench—58.3
PIQA88.7%—

Math QwQ-32B leads

GPT-4o mini: 10.4 (#314), QwQ-32B: 38.0 (#143)

Math benchmarks
BenchmarkGPT-4o miniQwQ-32B
OTIS Mock AIME 2024-20256.9%59.2%
LiveBench Math36.3%77.8%
LMArena Math12671359
FrontierMath (Tiers 1-3)0.7%—
Omni-MATH28%—
MATH Level 552.6%—
GSM8K91.3%—

Knowledge QwQ-32B leads

GPT-4o mini: 17.7 (#284), QwQ-32B: 37.2 (#158)

Knowledge benchmarks
BenchmarkGPT-4o miniQwQ-32B
GPQA Diamond37.7%65.3%
Confabulations37.2%15.6%
LMArena Expert12351324
SimpleQA Verified8.3%—
MMLU-Pro60.3%—
GPQA (HELM)36.8%—
BoolQ88.7%—
MMLU81.8%—

Multimodal Not comparable

GPT-4o mini: 25.9 (#122), QwQ-32B: —

Multimodal benchmarks
BenchmarkGPT-4o miniQwQ-32B
LMArena Vision1066—
Video-MME64.8%—
GeoBench64%—
VPCT34%—

Multilingual QwQ-32B leads

GPT-4o mini: 42.0 (#199), QwQ-32B: 44.8 (#176)

Multilingual benchmarks
BenchmarkGPT-4o miniQwQ-32B
LMArena Non-English12661305
LMArena Chinese12651378
LMArena French12971336
LMArena German12721313
LMArena Japanese12161262
LMArena Korean11951279
LMArena Russian12751297
LMArena Spanish12761354

Instruction Following QwQ-32B leads

GPT-4o mini: 61.9 (#239), QwQ-32B: 72.6 (#137)

Instruction Following benchmarks
BenchmarkGPT-4o miniQwQ-32B
LiveBench Instruction Following56.8%81.8%
LMArena Instruction Following12581297
IFEval78.2%—

Long Context QwQ-32B leads

GPT-4o mini: 39.1 (#186), QwQ-32B: 49.0 (#11)

Long Context benchmarks
BenchmarkGPT-4o miniQwQ-32B
LMArena Longer Query12891308
Fiction.LiveBench—83.3%

Writing & Preference QwQ-32B leads

GPT-4o mini: 39.5 (#248), QwQ-32B: 50.6 (#180)

Writing & Preference benchmarks
BenchmarkGPT-4o miniQwQ-32B
LMArena Text12861329
LMArena Creative Writing12681288
Short-Story Creative Writing67.2%80.2%
EQ-Bench Creative Writing8731257
LMArena Multi-Turn12851314
LiveBench Language28.6%51.4%
WildBench79.1%—

Frequently asked questions

Is GPT-4o mini better than QwQ-32B?

QwQ-32B is the stronger model overall, scoring 39.8 to 25.5 on the Noometry Index.

Is GPT-4o mini or QwQ-32B better for coding?

QwQ-32B scores higher on coding benchmarks: 35.4 versus 22.0 in the Noometry coding category.

How many benchmarks do GPT-4o mini and QwQ-32B share?

34 benchmarks have published results for both models. GPT-4o mini has 60 scored results on Noometry and QwQ-32B has 36.

Related comparisons

Go deeper