Microsoft, open weights

Phi 3 Small 8k Instruct

Phi 3 Small 8k Instruct by Microsoft ranks 314th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.3. Its strongest category is knowledge, where it ranks 240th.

Last verified

Specifications

Noometry rank
#314 of 354
Index score
29.3
Evidence
Confirmed 32 results
Provider
Microsoft
Released
April 23, 2024
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Phi 3 Small 8k Instruct category scores
  1. Coding 27.9
  2. Reasoning 14.8
  3. Math 27.6
  4. Knowledge 29.1
  5. Multilingual 28.5
  6. Instruction Following 51.9
  7. Long Context 33.0
  8. Writing & Preference 31.1
Phi 3 Small 8k Instruct category ranks
CategoryScoreRankResults
Coding27.9#3182
Reasoning14.8#3233
Math27.6#2482
Knowledge29.1#2401
Multilingual28.5#2721
Instruction Following51.9#2922
Long Context33.0#2671
Writing & Preference31.1#2844

Strengths and weaknesses

Categories where Phi 3 Small 8k Instruct places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Phi 3 Small 8k Instruct: strongest categories
CategoryScorevs medianRank
Math27.6−9.0#248 of 327, top 76%
Knowledge29.1−8.3#240 of 314, top 77%
Long Context33.0−7.9#267 of 296, top 91%

Weakest categories

Phi 3 Small 8k Instruct: weakest categories
CategoryScorevs medianRank
Instruction Following51.9−19.4#292 of 305, top 96%
Coding27.9−10.9#318 of 340, top 94%
Reasoning14.8−8.8#323 of 350, top 93%

Closest competitors

The models ranked just above and below Phi 3 Small 8k Instruct. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Phi 3 Small 8k Instruct
ModelRankScoreBlended $/MSpeed
Claude 3 Opus#31029.5——Compare
DBRX#31129.4——Compare
Gemma 2 27B#31229.4$0.65—Compare
Gemma 1.1 2b IT#31329.3——Compare
Claude 3.5 Haiku#31529.2——Compare
GPT-4#31629.1$37.50—Compare
Llama 2-7B#31729.1——Compare
Granite 4.0 Micro#31829.0$0.0408—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Phi 3 Small 8k Instruct Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Coding20.3%#34 of 39, top 88%Epoch AI
LMArena Coding1101#269 of 294, top 92%LMArena2026-10-08

Reasoning

Phi 3 Small 8k Instruct Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Reasoning15.9%#38 of 39, top 98%Epoch AI
LMArena Hard Prompts1100#268 of 297, top 91%LMArena2026-10-08
LiveBench Data Analysis30.3%#38 of 39, top 98%Epoch AI
Adversarial NLI58.1%#2 of 9, top 23%Epoch AI
BIG-Bench Hard79.1%#6 of 27, top 23%Epoch AI
HellaSwag77%#21 of 29, top 73%Epoch AI
LiveBench24%#37 of 39, top 95%Epoch AI
WinoGrande81.5%#13 of 43, top 31%Epoch AI

Math

Phi 3 Small 8k Instruct Math benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Math17.6%#37 of 39, top 95%Epoch AI
LMArena Math1151#252 of 285, top 89%LMArena2026-10-08

Knowledge

Phi 3 Small 8k Instruct Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Expert1067#257 of 273, top 95%LMArena2026-10-08
ARC (AI2) Challenge90.7%#6 of 39, top 16%Epoch AI
MMLU75.7%#33 of 81, top 41%Epoch AI
OpenBookQA88%#2 of 19, top 11%Epoch AI
TriviaQA58.1%#23 of 25, top 92%Epoch AI

Multilingual

Phi 3 Small 8k Instruct Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1058#272 of 297, top 92%LMArena2026-10-08
LMArena Chinese1061#264 of 285, top 93%LMArena2026-10-08
LMArena French1135#209 of 223, top 94%LMArena2026-10-08
LMArena German1080#215 of 231, top 94%LMArena2026-10-08
LMArena Japanese966#204 of 211, top 97%LMArena2026-10-08
LMArena Korean894#211 of 213, top 100%LMArena2026-10-08
LMArena Russian1111#255 of 283, top 91%LMArena2026-10-08
LMArena Spanish1111#213 of 226, top 95%LMArena2026-10-08

Instruction Following

Phi 3 Small 8k Instruct Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following47.2%#38 of 39, top 98%Epoch AI
LMArena Instruction Following1087#272 of 298, top 92%LMArena2026-10-08

Long Context

Phi 3 Small 8k Instruct Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1088#272 of 291, top 94%LMArena2026-10-08

Writing & Preference

Phi 3 Small 8k Instruct Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1110#270 of 297, top 91%LMArena2026-10-08
LMArena Creative Writing1083#272 of 295, top 93%LMArena2026-10-08
LMArena Multi-Turn1068#271 of 295, top 92%LMArena2026-10-08
LiveBench Language12.9%#37 of 39, top 95%Epoch AI

Compare Phi 3 Small 8k Instruct

Other Microsoft models

Frequently asked questions

How good is Phi 3 Small 8k Instruct?

Phi 3 Small 8k Instruct by Microsoft ranks 314th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.3. Its strongest category is knowledge, where it ranks 240th.

Is Phi 3 Small 8k Instruct open source?

Yes. Phi 3 Small 8k Instruct's weights are downloadable; check the license for commercial terms.

What are Phi 3 Small 8k Instruct's strengths and weaknesses?

Relative to other ranked models, Phi 3 Small 8k Instruct places best in math, knowledge, long context and lowest in instruction following, coding, reasoning.

What is Phi 3 Small 8k Instruct best at?

Its best category is knowledge, where it ranks 240th on Noometry.