Model comparison

Mistral Small 3.1 vs Phi-4 Mini

Mistral Small 3.1 and Phi-4 Mini score almost the same on the Noometry Index (31.7 vs 30.9), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Summary

  • The widest gap is in coding, where Mistral Small 3.1 leads 38.3 to 28.1.
  • Phi-4 Mini is cheaper at $0.075 / $0.30 per million input/output tokens, against $0.35 / $0.56 for Mistral Small 3.1.

Side by side

Mistral Small 3.1 and Phi-4 Mini specifications
Mistral Small 3.1Phi-4 Mini
ProviderMistral AIMicrosoft
Noometry Index31.730.9
Released2025-03-172024-12-11
WeightsOpenOpen
Context window128K128K
Max output102K4K
Input $ / M tokens$0.35$0.075
Output $ / M tokens$0.56$0.30
Results tracked283

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3.1 leads

Mistral Small 3.1: 38.3 (#179), Phi-4 Mini: 28.1 (#317)

Coding benchmarks
BenchmarkMistral Small 3.1Phi-4 Mini
SciCode—10.8%
LMArena Coding1309—

Reasoning Phi-4 Mini leads

Mistral Small 3.1: 19.7 (#254), Phi-4 Mini: 22.4 (#195)

Reasoning benchmarks
BenchmarkMistral Small 3.1Phi-4 Mini
CritPt—0%
Chess Puzzles1%—
LMArena Hard Prompts1278—
Epoch Capabilities Index127.48—

Math Not comparable

Mistral Small 3.1: 14.7 (#301), Phi-4 Mini: —

Math benchmarks
BenchmarkMistral Small 3.1Phi-4 Mini
OTIS Mock AIME 2024-20253.9%—
Omni-MATH24.8%—
LMArena Math1262—

Knowledge Phi-4 Mini leads

Mistral Small 3.1: 22.6 (#271), Phi-4 Mini: 25.3 (#262)

Knowledge benchmarks
BenchmarkMistral Small 3.1Phi-4 Mini
GPQA Diamond41.9%—
MMLU-Pro61%—
Vectara Hallucination Rate—23.5%
GPQA (HELM)39.2%—
LMArena Expert1257—

Multimodal Not comparable

Mistral Small 3.1: 33.2 (#99), Phi-4 Mini: —

Multimodal benchmarks
BenchmarkMistral Small 3.1Phi-4 Mini
LMArena Vision1136—

Multilingual Not comparable

Mistral Small 3.1: 41.2 (#209), Phi-4 Mini: —

Multilingual benchmarks
BenchmarkMistral Small 3.1Phi-4 Mini
LMArena Non-English1255—
LMArena Chinese1253—
LMArena French1273—
LMArena German1266—
LMArena Japanese1208—
LMArena Korean1206—
LMArena Russian1263—
LMArena Spanish1283—

Instruction Following Not comparable

Mistral Small 3.1: 63.6 (#230), Phi-4 Mini: —

Instruction Following benchmarks
BenchmarkMistral Small 3.1Phi-4 Mini
IFEval75%—
LMArena Instruction Following1264—

Long Context Not comparable

Mistral Small 3.1: 39.5 (#178), Phi-4 Mini: —

Long Context benchmarks
BenchmarkMistral Small 3.1Phi-4 Mini
LMArena Longer Query1299—

Writing & Preference Not comparable

Mistral Small 3.1: 37.0 (#259), Phi-4 Mini: —

Writing & Preference benchmarks
BenchmarkMistral Small 3.1Phi-4 Mini
LMArena Text1277—
LMArena Creative Writing1253—
EQ-Bench Creative Writing761—
WildBench78.8%—
LMArena Multi-Turn1270—

Frequently asked questions

Is Mistral Small 3.1 better than Phi-4 Mini?

Mistral Small 3.1 and Phi-4 Mini score almost the same on the Noometry Index (31.7 vs 30.9), so choose on price, context window or the category you care about most.

Which is cheaper, Mistral Small 3.1 or Phi-4 Mini?

Phi-4 Mini is cheaper. It lists at $0.075 per million input tokens and $0.30 per million output tokens; Mistral Small 3.1 lists at $0.35 and $0.56.

Is Mistral Small 3.1 or Phi-4 Mini better for coding?

Mistral Small 3.1 scores higher on coding benchmarks: 38.3 versus 28.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Mistral Small 3.1 and Phi-4 Mini share?

0 benchmarks have published results for both models. Mistral Small 3.1 has 28 scored results on Noometry and Phi-4 Mini has 3.

Related comparisons

Go deeper