Model comparison

Claude 3.5 Sonnet vs Phi-4

Claude 3.5 Sonnet is the stronger model overall, scoring 34.6 to 31.2 on the Noometry Index.

Last verified . 34 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • They share 34 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 7 categories and Phi-4 in 2 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Claude 3.5 Sonnet leads 52.9 to 40.5.
  • The biggest single-benchmark swing is LiveBench Coding: 67.1% for Claude 3.5 Sonnet and 30.7% for Phi-4.
  • Phi-4 has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Phi-4 specifications
Claude 3.5 SonnetPhi-4
ProviderAnthropicMicrosoft
Noometry Index34.631.2
Released2024-06-202024-12-11
WeightsProprietaryOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.07
Output $ / M tokens—$0.14
Results tracked6037

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkClaude 3.5 SonnetPhi-4
BigCodeBench Instruct46.8%45.5%
LiveBench Coding67.1%30.7%
LMArena Coding13421231
BigCodeBench Complete58.6%55.4%
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 32.3 (#67), Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetPhi-4
BALROG32.6%11.6%
Berkeley Function Calling Leaderboard—28.8%
TheAgentCompany24%—
Cybench17.5%—
METR Time Horizons45.2%—

Reasoning Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 23.1 (#183), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetPhi-4
LiveBench Reasoning56.7%47.8%
LMArena Hard Prompts13051220
LiveBench Data Analysis55%45.2%
Epoch Capabilities Index133.55130.42
LiveBench59%41.6%
SimpleBench41.4%—
Chess Puzzles—1%
EnigmaEval0.9%—
DTBench67.8%—
ForecastBench60.7—

Math Phi-4 leads

Claude 3.5 Sonnet: 19.2 (#288), Phi-4: 20.8 (#285)

Math benchmarks
BenchmarkClaude 3.5 SonnetPhi-4
OTIS Mock AIME 2024-20258.5%13.8%
LiveBench Math52.3%42%
LMArena Math13071246
MATH Level 556.9%64.9%
Omni-MATH27.6%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Phi-4 leads

Claude 3.5 Sonnet: 28.6 (#245), Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetPhi-4
GPQA Diamond55.3%56.1%
Confabulations19.9%29.4%
LMArena Expert12651203
MMLU87.3%84.8%
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Vectara Hallucination Rate—3.7%
GPQA (HELM)56.5%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Phi-4: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetPhi-4
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 43.2 (#185), Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetPhi-4
LMArena Non-English12831197
LMArena Chinese12721212
LMArena French13051224
LMArena German12971222
LMArena Japanese12341158
LMArena Korean12001151
LMArena Russian13061209
LMArena Spanish12901234

Instruction Following Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 68.8 (#182), Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetPhi-4
LiveBench Instruction Following69.3%58.4%
LMArena Instruction Following12971201
IFEval85.5%—

Long Context Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.9 (#167), Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetPhi-4
LMArena Longer Query13111217

Writing & Preference Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 52.9 (#164), Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetPhi-4
LMArena Text12981217
LMArena Creative Writing12921182
Short-Story Creative Writing80.3%62.6%
LMArena Multi-Turn13261206
LiveBench Language53.8%25.6%
EQ-Bench Creative Writing1451—
WildBench79.2%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Phi-4?

Claude 3.5 Sonnet is the stronger model overall, scoring 34.6 to 31.2 on the Noometry Index.

Is Claude 3.5 Sonnet or Phi-4 better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 34.4 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Phi-4 share?

34 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper