Model comparison

Qwen2.5-Coder-32B vs Solar Pro4

Solar Pro4 is the stronger model overall, scoring 42.1 to 33.4 on the Noometry Index.

Last verified . 12 shared benchmarks.

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Solar Pro4 Upstage

42.1

Rank #121 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Qwen2.5-Coder-32B scores higher in 0 categories and Solar Pro4 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Solar Pro4 leads 40.1 to 22.6.
  • Solar Pro4 is cheaper at $0.30 / $1.20 per million input/output tokens, against $0.66 / $1 for Qwen2.5-Coder-32B.
  • Solar Pro4 accepts more context: 524K tokens versus 33K.
  • Qwen2.5-Coder-32B has downloadable open weights; the other is API-only.

Side by side

Qwen2.5-Coder-32B and Solar Pro4 specifications
Qwen2.5-Coder-32BSolar Pro4
ProviderAlibaba (Qwen)Upstage
Noometry Index33.442.1
Released2024-09-182026-08-06
WeightsOpenProprietary
Context window33K524K
Max output29K131K
Input $ / M tokens$0.66$0.30
Output $ / M tokens$1$1.20
Results tracked3118

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Solar Pro4 leads

Qwen2.5-Coder-32B: 22.6 (#333), Solar Pro4: 40.1 (#149)

Coding benchmarks
BenchmarkQwen2.5-Coder-32BSolar Pro4
LMArena Coding12761437
SWE-bench Verified (bash only)9%—
Aider Polyglot16.4%—
LMArena WebDev—1371
BigCodeBench Instruct49%—
LiveBench Coding56.9%—
BigCodeBench Complete58%—
HumanEval+87.2%—
MBPP+77%—

Reasoning Solar Pro4 leads

Qwen2.5-Coder-32B: 21.2 (#225), Solar Pro4: 28.5 (#104)

Reasoning benchmarks
BenchmarkQwen2.5-Coder-32BSolar Pro4
LMArena Hard Prompts12511399
LiveBench Reasoning42.1%—
LiveBench Data Analysis49.9%—
Epoch Capabilities Index119.49—
HellaSwag83%—
LiveBench46.2%—
WinoGrande80.8%—

Math Solar Pro4 leads

Qwen2.5-Coder-32B: 33.3 (#204), Solar Pro4: 38.8 (#128)

Math benchmarks
BenchmarkQwen2.5-Coder-32BSolar Pro4
LMArena Math12511416
LiveBench Math46.6%—
GSM8K93%—

Knowledge Solar Pro4 leads

Qwen2.5-Coder-32B: 33.4 (#203), Solar Pro4: 39.8 (#129)

Knowledge benchmarks
BenchmarkQwen2.5-Coder-32BSolar Pro4
LMArena Expert12211427
ARC (AI2) Challenge70.5%—
MMLU79.1%—

Multilingual Solar Pro4 leads

Qwen2.5-Coder-32B: 37.8 (#235), Solar Pro4: 48.7 (#139)

Multilingual benchmarks
BenchmarkQwen2.5-Coder-32BSolar Pro4
LMArena Non-English12051361
LMArena Chinese12221415
LMArena Russian12281360
LMArena French—1397
LMArena German—1364
LMArena Japanese—1309
LMArena Korean—1382
LMArena Spanish—1401

Instruction Following Solar Pro4 leads

Qwen2.5-Coder-32B: 61.4 (#245), Solar Pro4: 72.7 (#132)

Instruction Following benchmarks
BenchmarkQwen2.5-Coder-32BSolar Pro4
LMArena Instruction Following12231377
LiveBench Instruction Following58.7%—

Long Context Solar Pro4 leads

Qwen2.5-Coder-32B: 38.0 (#208), Solar Pro4: 42.1 (#130)

Long Context benchmarks
BenchmarkQwen2.5-Coder-32BSolar Pro4
LMArena Longer Query12511381

Writing & Preference Solar Pro4 leads

Qwen2.5-Coder-32B: 41.6 (#240), Solar Pro4: 56.5 (#138)

Writing & Preference benchmarks
BenchmarkQwen2.5-Coder-32BSolar Pro4
LMArena Text12301386
LMArena Creative Writing11741316
LMArena Multi-Turn12221385
LiveBench Language23.3%—

Frequently asked questions

Is Qwen2.5-Coder-32B better than Solar Pro4?

Solar Pro4 is the stronger model overall, scoring 42.1 to 33.4 on the Noometry Index.

Which is cheaper, Qwen2.5-Coder-32B or Solar Pro4?

Solar Pro4 is cheaper. It lists at $0.30 per million input tokens and $1.20 per million output tokens; Qwen2.5-Coder-32B lists at $0.66 and $1.

Is Qwen2.5-Coder-32B or Solar Pro4 better for coding?

Solar Pro4 scores higher on coding benchmarks: 40.1 versus 22.6 in the Noometry coding category.

Which has the bigger context window?

Solar Pro4 does, with 524K tokens against 33K.

How many benchmarks do Qwen2.5-Coder-32B and Solar Pro4 share?

12 benchmarks have published results for both models. Qwen2.5-Coder-32B has 31 scored results on Noometry and Solar Pro4 has 18.

Related comparisons

Go deeper