Model comparison

Claude 3 Sonnet vs Phi 3 Mini 128k Instruct

Claude 3 Sonnet and Phi 3 Mini 128k Instruct score almost the same on the Noometry Index (29.0 vs 29.7), so choose on price, context window or the category you care about most.

Last verified . 19 shared benchmarks.

Claude 3 Sonnet Anthropic

29.0

Rank #319 Confirmed

Phi 3 Mini 128k Instruct Microsoft

29.7

Rank #305 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Claude 3 Sonnet scores higher in 6 categories and Phi 3 Mini 128k Instruct in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Phi 3 Mini 128k Instruct leads 31.6 to 10.7.
  • The biggest single-benchmark swing is BigCodeBench Complete: 53.8% for Claude 3 Sonnet and 40.6% for Phi 3 Mini 128k Instruct.
  • Phi 3 Mini 128k Instruct has downloadable open weights; the other is API-only.

Side by side

Claude 3 Sonnet and Phi 3 Mini 128k Instruct specifications
Claude 3 SonnetPhi 3 Mini 128k Instruct
ProviderAnthropicMicrosoft
Noometry Index29.029.7
Released2024-02-292024-04-23
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3019

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3 Sonnet: 29.6 (#302), Phi 3 Mini 128k Instruct: 28.8 (#312)

Coding benchmarks
BenchmarkClaude 3 SonnetPhi 3 Mini 128k Instruct
BigCodeBench Instruct42.7%29.6%
LMArena Coding12231039
BigCodeBench Complete53.8%40.6%
WeirdML10.2%—
HumanEval+64%—
MBPP+69.3%—

Reasoning Too close to call

Claude 3 Sonnet: 20.5 (#237), Phi 3 Mini 128k Instruct: 19.6 (#256)

Reasoning benchmarks
BenchmarkClaude 3 SonnetPhi 3 Mini 128k Instruct
LMArena Hard Prompts11971028
DTBench53.6%—
Epoch Capabilities Index120.7—
WinoGrande75.1%—

Math Phi 3 Mini 128k Instruct leads

Claude 3 Sonnet: 10.7 (#310), Phi 3 Mini 128k Instruct: 31.6 (#222)

Math benchmarks
BenchmarkClaude 3 SonnetPhi 3 Mini 128k Instruct
LMArena Math12131089
OTIS Mock AIME 2024-20252.5%—
MATH Level 518.2%—

Knowledge Phi 3 Mini 128k Instruct leads

Claude 3 Sonnet: 21.1 (#276), Phi 3 Mini 128k Instruct: 26.8 (#254)

Knowledge benchmarks
BenchmarkClaude 3 SonnetPhi 3 Mini 128k Instruct
LMArena Expert1173984
GPQA Diamond40.6%—
MMLU75.9%—

Multimodal Not comparable

Claude 3 Sonnet: 25.2 (#125), Phi 3 Mini 128k Instruct: —

Multimodal benchmarks
BenchmarkClaude 3 SonnetPhi 3 Mini 128k Instruct
LMArena Vision984—

Multilingual Claude 3 Sonnet leads

Claude 3 Sonnet: 37.8 (#234), Phi 3 Mini 128k Instruct: 25.2 (#285)

Multilingual benchmarks
BenchmarkClaude 3 SonnetPhi 3 Mini 128k Instruct
LMArena Non-English12051000
LMArena Chinese11891016
LMArena French12291039
LMArena German12041006
LMArena Japanese1131899
LMArena Korean1128856
LMArena Russian12271004
LMArena Spanish12041059

Instruction Following Claude 3 Sonnet leads

Claude 3 Sonnet: 62.8 (#235), Phi 3 Mini 128k Instruct: 51.8 (#294)

Instruction Following benchmarks
BenchmarkClaude 3 SonnetPhi 3 Mini 128k Instruct
LMArena Instruction Following11991023

Long Context Claude 3 Sonnet leads

Claude 3 Sonnet: 36.7 (#228), Phi 3 Mini 128k Instruct: 30.4 (#289)

Long Context benchmarks
BenchmarkClaude 3 SonnetPhi 3 Mini 128k Instruct
LMArena Longer Query1211996

Writing & Preference Claude 3 Sonnet leads

Claude 3 Sonnet: 42.1 (#238), Phi 3 Mini 128k Instruct: 27.1 (#301)

Writing & Preference benchmarks
BenchmarkClaude 3 SonnetPhi 3 Mini 128k Instruct
LMArena Text12181050
LMArena Creative Writing11861024
LMArena Multi-Turn1227989

Frequently asked questions

Is Claude 3 Sonnet better than Phi 3 Mini 128k Instruct?

Claude 3 Sonnet and Phi 3 Mini 128k Instruct score almost the same on the Noometry Index (29.0 vs 29.7), so choose on price, context window or the category you care about most.

Is Claude 3 Sonnet or Phi 3 Mini 128k Instruct better for coding?

They score almost the same on coding (29.6 vs 28.8); test both on your own repository before choosing.

How many benchmarks do Claude 3 Sonnet and Phi 3 Mini 128k Instruct share?

19 benchmarks have published results for both models. Claude 3 Sonnet has 30 scored results on Noometry and Phi 3 Mini 128k Instruct has 19.

Related comparisons

Go deeper