Model comparison

MiniMax-M2.5 vs Qwen1.5-110B

MiniMax-M2.5 is the stronger model overall, scoring 38.3 to 34.2 on the Noometry Index.

Last verified . 17 shared benchmarks.

MiniMax-M2.5 MiniMax

38.3

Rank #188 Confirmed

Qwen1.5-110B Alibaba (Qwen)

34.2

Rank #234 Confirmed

Summary

  • They share 17 benchmarks with published results for both. MiniMax-M2.5 scores higher in 6 categories and Qwen1.5-110B in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where MiniMax-M2.5 leads 53.9 to 38.0.

Side by side

MiniMax-M2.5 and Qwen1.5-110B specifications
MiniMax-M2.5Qwen1.5-110B
ProviderMiniMaxAlibaba (Qwen)
Noometry Index38.334.2
Released2026-02-122024-04-25
WeightsOpenOpen
Context window205K—
Max output131K—
Input $ / M tokens$0.30—
Output $ / M tokens$1.20—
Results tracked3320

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.5 leads

MiniMax-M2.5: 48.1 (#58), Qwen1.5-110B: 33.0 (#264)

Coding benchmarks
BenchmarkMiniMax-M2.5Qwen1.5-110B
LMArena Coding13811184
SWE-bench Verified (bash only)75.8%—
LMArena WebDev1387—
SWE-bench Multilingual68.3%—
BigCodeBench Instruct—35%
BigCodeBench Complete—44.4%
ALE-Bench618.17—

Agentic & Tool Use Not comparable

MiniMax-M2.5: 30.4 (#77), Qwen1.5-110B: —

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M2.5Qwen1.5-110B
Terminal-Bench42.7%—
Vending-Bench 2-23.16—

Reasoning Qwen1.5-110B leads

MiniMax-M2.5: 17.5 (#292), Qwen1.5-110B: 22.7 (#189)

Reasoning benchmarks
BenchmarkMiniMax-M2.5Qwen1.5-110B
LMArena Hard Prompts13721168
ARC-AGI-24.9%—
Kagi LLM Benchmark55.2%—
NYT Connections (extended)16.8%—
ARC-AGI-163.7%—
Epoch Capabilities Index146.68—
ForecastBench—57.7

Math Qwen1.5-110B leads

MiniMax-M2.5: 26.9 (#253), Qwen1.5-110B: 33.7 (#201)

Math benchmarks
BenchmarkMiniMax-M2.5Qwen1.5-110B
LMArena Math13781185
ProofBench4%—

Knowledge MiniMax-M2.5 leads

MiniMax-M2.5: 39.2 (#135), Qwen1.5-110B: 31.2 (#219)

Knowledge benchmarks
BenchmarkMiniMax-M2.5Qwen1.5-110B
LMArena Expert13791144
Vectara Hallucination Rate9.1%—

Multilingual MiniMax-M2.5 leads

MiniMax-M2.5: 47.1 (#152), Qwen1.5-110B: 33.6 (#250)

Multilingual benchmarks
BenchmarkMiniMax-M2.5Qwen1.5-110B
LMArena Non-English13381142
LMArena Chinese13931206
LMArena French13621151
LMArena German13621123
LMArena Japanese11711074
LMArena Korean12321044
LMArena Russian13581118
LMArena Spanish13541142

Instruction Following MiniMax-M2.5 leads

MiniMax-M2.5: 71.5 (#148), Qwen1.5-110B: 60.3 (#252)

Instruction Following benchmarks
BenchmarkMiniMax-M2.5Qwen1.5-110B
LMArena Instruction Following13531158

Long Context MiniMax-M2.5 leads

MiniMax-M2.5: 37.5 (#216), Qwen1.5-110B: 35.1 (#242)

Long Context benchmarks
BenchmarkMiniMax-M2.5Qwen1.5-110B
LMArena Longer Query13661157
CL-bench11.4%—
CL-bench Life6.3%—

Writing & Preference MiniMax-M2.5 leads

MiniMax-M2.5: 53.9 (#153), Qwen1.5-110B: 38.0 (#255)

Writing & Preference benchmarks
BenchmarkMiniMax-M2.5Qwen1.5-110B
LMArena Text13591175
LMArena Creative Writing13311148
LMArena Multi-Turn13641160
EQ-Bench Creative Writing1361—

Frequently asked questions

Is MiniMax-M2.5 better than Qwen1.5-110B?

MiniMax-M2.5 is the stronger model overall, scoring 38.3 to 34.2 on the Noometry Index.

Is MiniMax-M2.5 or Qwen1.5-110B better for coding?

MiniMax-M2.5 scores higher on coding benchmarks: 48.1 versus 33.0 in the Noometry coding category.

How many benchmarks do MiniMax-M2.5 and Qwen1.5-110B share?

17 benchmarks have published results for both models. MiniMax-M2.5 has 33 scored results on Noometry and Qwen1.5-110B has 20.

Related comparisons

Go deeper