Model comparison
Granite 4.0 Micro vs phi-3-medium 14B
Granite 4.0 Micro and phi-3-medium 14B score almost the same on the Noometry Index (29.0 vs 29.7), so choose on price, context window or the category you care about most.
Last verified . 1 shared benchmarks.
Summary
- They share 1 benchmark with published results for both. Granite 4.0 Micro scores higher in 1 category and phi-3-medium 14B in 1 category; one gap is clear of the uncertainty.
- The widest gap is in math, where phi-3-medium 14B leads 27.3 to 12.0.
Side by side
| Granite 4.0 Micro | phi-3-medium 14B | |
|---|---|---|
| Provider | IBM | Microsoft |
| Noometry Index | 29.0 | 29.7 |
| Released | 2025-10-02 | 2024-04-23 |
| Weights | Open | Open |
| Context window | 131K | — |
| Max output | 118K | — |
| Input $ / M tokens | $0.017 | — |
| Output $ / M tokens | $0.11 | — |
| Results tracked | 8 | 13 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Granite 4.0 Micro: —, phi-3-medium 14B: 36.8 (#201)
| Benchmark | Granite 4.0 Micro | phi-3-medium 14B |
|---|---|---|
| BigCodeBench Instruct | — | 37.6% |
| BigCodeBench Complete | — | 48.7% |
Reasoning Not comparable
Granite 4.0 Micro: 19.2 (#265), phi-3-medium 14B: —
| Benchmark | Granite 4.0 Micro | phi-3-medium 14B |
|---|---|---|
| Chess Puzzles | 0% | — |
| Adversarial NLI | — | 55.8% |
| BIG-Bench Hard | — | 81.4% |
| Epoch Capabilities Index | — | 121.23 |
| HellaSwag | — | 82.4% |
| WinoGrande | — | 81.5% |
Math phi-3-medium 14B leads
Granite 4.0 Micro: 12.0 (#307), phi-3-medium 14B: 27.3 (#250)
| Benchmark | Granite 4.0 Micro | phi-3-medium 14B |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 2.8% | — |
| Omni-MATH | 20.9% | — |
| MATH Level 5 | — | 17.6% |
Knowledge Too close to call
Granite 4.0 Micro: 9.9 (#304), phi-3-medium 14B: 9.1 (#306)
| Benchmark | Granite 4.0 Micro | phi-3-medium 14B |
|---|---|---|
| GPQA Diamond | 28.3% | 27.6% |
| MMLU-Pro | 39.5% | — |
| GPQA (HELM) | 30.7% | — |
| ARC (AI2) Challenge | — | 91.6% |
| MMLU | — | 78% |
| OpenBookQA | — | 87.4% |
| TriviaQA | — | 73.9% |
Instruction Following Not comparable
Granite 4.0 Micro: 69.9 (#169), phi-3-medium 14B: —
| Benchmark | Granite 4.0 Micro | phi-3-medium 14B |
|---|---|---|
| IFEval | 84.9% | — |
Writing & Preference Not comparable
Granite 4.0 Micro: 46.7 (#216), phi-3-medium 14B: —
| Benchmark | Granite 4.0 Micro | phi-3-medium 14B |
|---|---|---|
| WildBench | 67% | — |
Frequently asked questions
Is Granite 4.0 Micro better than phi-3-medium 14B?
Granite 4.0 Micro and phi-3-medium 14B score almost the same on the Noometry Index (29.0 vs 29.7), so choose on price, context window or the category you care about most.
How many benchmarks do Granite 4.0 Micro and phi-3-medium 14B share?
1 benchmark has published results for both models. Granite 4.0 Micro has 8 scored results on Noometry and phi-3-medium 14B has 13.