Model comparison

Mistral Medium 3.5 vs Phi-4

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 31.2 on the Noometry Index. Phi-4 costs 34× less per token, which makes it the better buy when Mistral Medium 3.5's lead doesn't matter for your workload.

Last verified . 17 shared benchmarks.

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 7 categories and Phi-4 in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Medium 3.5 leads 39.1 to 20.8.
  • Phi-4 is cheaper at $0.07 / $0.14 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium 3.5.
  • Mistral Medium 3.5 accepts more context: 262K tokens versus 128K.

Side by side

Mistral Medium 3.5 and Phi-4 specifications
Mistral Medium 3.5Phi-4
ProviderMistral AIMicrosoft
Noometry Index40.231.2
Released—2024-12-11
WeightsOpenOpen
Context window262K128K
Max output210K4K
Input $ / M tokens$1.50$0.07
Output $ / M tokens$7.50$0.14
Results tracked2237

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium 3.5 leads

Mistral Medium 3.5: 36.0 (#213), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkMistral Medium 3.5Phi-4
LMArena Coding14611231
LMArena WebDev1264—
BigCodeBench Instruct—45.5%
LiveBench Coding—30.7%
BigCodeBench Complete—55.4%

Agentic & Tool Use Not comparable

Mistral Medium 3.5: —, Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkMistral Medium 3.5Phi-4
Berkeley Function Calling Leaderboard—28.8%
BALROG—11.6%

Reasoning Too close to call

Mistral Medium 3.5: 17.3 (#295), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkMistral Medium 3.5Phi-4
LMArena Hard Prompts14361220
Epoch Capabilities Index141.35130.42
Kagi LLM Benchmark41.4%—
NYT Connections (extended)12.9%—
Chess Puzzles—1%
LiveBench Reasoning—47.8%
LiveBench Data Analysis—45.2%
LiveBench—41.6%

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), Phi-4: 20.8 (#285)

Math benchmarks
BenchmarkMistral Medium 3.5Phi-4
LMArena Math14311246
OTIS Mock AIME 2024-2025—13.8%
LiveBench Math—42%
MATH Level 5—64.9%

Knowledge Mistral Medium 3.5 leads

Mistral Medium 3.5: 40.0 (#126), Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkMistral Medium 3.5Phi-4
LMArena Expert14321203
GPQA Diamond—56.1%
Confabulations—29.4%
Vectara Hallucination Rate—3.7%
MMLU—84.8%

Multimodal Not comparable

Mistral Medium 3.5: 38.3 (#65), Phi-4: —

Multimodal benchmarks
BenchmarkMistral Medium 3.5Phi-4
LMArena Vision1223—

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkMistral Medium 3.5Phi-4
LMArena Non-English14041197
LMArena Chinese14421212
LMArena French14481224
LMArena German14511222
LMArena Korean13851151
LMArena Russian13951209
LMArena Spanish14091234
LMArena Japanese—1158

Instruction Following Mistral Medium 3.5 leads

Mistral Medium 3.5: 74.6 (#90), Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5Phi-4
LMArena Instruction Following14151201
LiveBench Instruction Following—58.4%

Long Context Mistral Medium 3.5 leads

Mistral Medium 3.5: 43.2 (#103), Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkMistral Medium 3.5Phi-4
LMArena Longer Query14151217

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5Phi-4
LMArena Text14211217
LMArena Creative Writing13741182
LMArena Multi-Turn14231206
Short-Story Creative Writing—62.6%
EQ-Bench 4993—
LiveBench Language—25.6%

Frequently asked questions

Is Mistral Medium 3.5 better than Phi-4?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 31.2 on the Noometry Index. Phi-4 costs 34× less per token, which makes it the better buy when Mistral Medium 3.5's lead doesn't matter for your workload.

Which is cheaper, Mistral Medium 3.5 or Phi-4?

Phi-4 is cheaper. It lists at $0.07 per million input tokens and $0.14 per million output tokens; Mistral Medium 3.5 lists at $1.50 and $7.50.

Is Mistral Medium 3.5 or Phi-4 better for coding?

Mistral Medium 3.5 scores higher on coding benchmarks: 36.0 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium 3.5 does, with 262K tokens against 128K.

How many benchmarks do Mistral Medium 3.5 and Phi-4 share?

17 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper