Model comparison

MiniMax M1 vs phi-3-medium 14B

MiniMax M1 is the stronger model overall, scoring 40.3 to 29.7 on the Noometry Index.

Last verified . 0 shared benchmarks.

MiniMax M1 MiniMax

40.3

Rank #150 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • The widest gap is in knowledge, where MiniMax M1 leads 36.4 to 9.1.

Side by side

MiniMax M1 and phi-3-medium 14B specifications
MiniMax M1phi-3-medium 14B
ProviderMiniMaxMicrosoft
Noometry Index40.329.7
Released2025-06-132024-04-23
WeightsOpenOpen
Context window1M—
Max output40K—
Input $ / M tokens$0.55—
Output $ / M tokens$2.20—
Results tracked1813

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax M1 leads

MiniMax M1: 39.9 (#153), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkMiniMax M1phi-3-medium 14B
BigCodeBench Instruct—37.6%
LMArena Coding1359—
BigCodeBench Complete—48.7%

Reasoning Not comparable

MiniMax M1: 26.9 (#126), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkMiniMax M1phi-3-medium 14B
LMArena Hard Prompts1339—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
Epoch Capabilities Index—121.23
HellaSwag—82.4%
WinoGrande—81.5%

Math MiniMax M1 leads

MiniMax M1: 37.5 (#151), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkMiniMax M1phi-3-medium 14B
LMArena Math1361—
MATH Level 5—17.6%

Knowledge MiniMax M1 leads

MiniMax M1: 36.4 (#170), phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkMiniMax M1phi-3-medium 14B
GPQA Diamond—27.6%
LMArena Expert1317—
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Multilingual Not comparable

MiniMax M1: 45.8 (#163), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkMiniMax M1phi-3-medium 14B
LMArena Non-English1319—
LMArena Chinese1360—
LMArena French1370—
LMArena German1350—
LMArena Japanese1217—
LMArena Korean1266—
LMArena Russian1329—
LMArena Spanish1353—

Instruction Following Not comparable

MiniMax M1: 69.3 (#174), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkMiniMax M1phi-3-medium 14B
LMArena Instruction Following1312—

Long Context Not comparable

MiniMax M1: 41.4 (#141), phi-3-medium 14B: —

Long Context benchmarks
BenchmarkMiniMax M1phi-3-medium 14B
Fiction.LiveBench69.4%—
LMArena Longer Query1326—

Writing & Preference Not comparable

MiniMax M1: 53.1 (#161), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkMiniMax M1phi-3-medium 14B
LMArena Text1343—
LMArena Creative Writing1298—
LMArena Multi-Turn1335—

Frequently asked questions

Is MiniMax M1 better than phi-3-medium 14B?

MiniMax M1 is the stronger model overall, scoring 40.3 to 29.7 on the Noometry Index.

Is MiniMax M1 or phi-3-medium 14B better for coding?

MiniMax M1 scores higher on coding benchmarks: 39.9 versus 36.8 in the Noometry coding category.

How many benchmarks do MiniMax M1 and phi-3-medium 14B share?

0 benchmarks have published results for both models. MiniMax M1 has 18 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper