Model comparison

MiniMax-M2 vs Qwen-14B

MiniMax-M2 is the stronger model overall, scoring 37.4 to 31.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

MiniMax-M2 MiniMax

37.4

Rank #204 Confirmed

Qwen-14B Alibaba (Qwen)

31.4

Rank #275 Confirmed

Summary

  • They share 10 benchmarks with published results for both. MiniMax-M2 scores higher in 6 categories and Qwen-14B in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where MiniMax-M2 leads 53.0 to 27.6.

Side by side

MiniMax-M2 and Qwen-14B specifications
MiniMax-M2Qwen-14B
ProviderMiniMaxAlibaba (Qwen)
Noometry Index37.431.4
Released2025-10-272023-09-24
WeightsOpenOpen
Context window205K—
Max output131K—
Input $ / M tokens$0.30—
Output $ / M tokens$1.20—
Results tracked2118

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2 leads

MiniMax-M2: 39.3 (#159), Qwen-14B: 31.2 (#288)

Coding benchmarks
BenchmarkMiniMax-M2Qwen-14B
LMArena Coding13701071
SWE-bench Verified (bash only)61%—
LMArena WebDev1297—

Agentic & Tool Use Not comparable

MiniMax-M2: 25.1 (#109), Qwen-14B: —

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M2Qwen-14B
Terminal-Bench30%—
Vending-Bench 2160.6—

Reasoning Too close to call

MiniMax-M2: 19.4 (#258), Qwen-14B: 19.6 (#257)

Reasoning benchmarks
BenchmarkMiniMax-M2Qwen-14B
LMArena Hard Prompts13571027
Kagi LLM Benchmark57.8%—
NYT Connections (extended)14.8%—
BIG-Bench Hard—55%
Epoch Capabilities Index—113.03
LAMBADA—71.1%
PIQA—79.9%

Math MiniMax-M2 leads

MiniMax-M2: 37.3 (#160), Qwen-14B: 31.2 (#227)

Math benchmarks
BenchmarkMiniMax-M2Qwen-14B
LMArena Math13521068
GSM8K—61.3%

Knowledge Not comparable

MiniMax-M2: 37.0 (#163), Qwen-14B: —

Knowledge benchmarks
BenchmarkMiniMax-M2Qwen-14B
LMArena Expert1337—
ARC (AI2) Challenge—84.4%
BoolQ—86.2%
MMLU—66.3%

Multilingual MiniMax-M2 leads

MiniMax-M2: 45.3 (#171), Qwen-14B: 27.5 (#275)

Multilingual benchmarks
BenchmarkMiniMax-M2Qwen-14B
LMArena Non-English13131041
LMArena Chinese13661077
LMArena French1335—
LMArena German1355—
LMArena Russian1331—
LMArena Spanish1326—

Instruction Following MiniMax-M2 leads

MiniMax-M2: 70.2 (#166), Qwen-14B: 52.4 (#289)

Instruction Following benchmarks
BenchmarkMiniMax-M2Qwen-14B
LMArena Instruction Following13281031

Long Context MiniMax-M2 leads

MiniMax-M2: 40.5 (#153), Qwen-14B: 31.3 (#280)

Long Context benchmarks
BenchmarkMiniMax-M2Qwen-14B
LMArena Longer Query13311028

Writing & Preference MiniMax-M2 leads

MiniMax-M2: 53.0 (#162), Qwen-14B: 27.6 (#299)

Writing & Preference benchmarks
BenchmarkMiniMax-M2Qwen-14B
LMArena Text13401051
LMArena Creative Writing12861028
LMArena Multi-Turn13611022

Frequently asked questions

Is MiniMax-M2 better than Qwen-14B?

MiniMax-M2 is the stronger model overall, scoring 37.4 to 31.4 on the Noometry Index.

Is MiniMax-M2 or Qwen-14B better for coding?

MiniMax-M2 scores higher on coding benchmarks: 39.3 versus 31.2 in the Noometry coding category.

How many benchmarks do MiniMax-M2 and Qwen-14B share?

10 benchmarks have published results for both models. MiniMax-M2 has 21 scored results on Noometry and Qwen-14B has 18.

Related comparisons

Go deeper