Model comparison

Llama 4 Maverick vs Phi 3 Mini 4k Instruct

Llama 4 Maverick is the stronger model overall, scoring 30.9 to 27.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Phi 3 Mini 4k Instruct Microsoft

27.9

Rank #328 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Llama 4 Maverick scores higher in 4 categories and Phi 3 Mini 4k Instruct in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama 4 Maverick leads 71.7 to 47.7.

Side by side

Llama 4 Maverick and Phi 3 Mini 4k Instruct specifications
Llama 4 MaverickPhi 3 Mini 4k Instruct
ProviderMetaMicrosoft
Noometry Index30.927.9
Released2025-04-052024-04-23
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.19—
Output $ / M tokens$0.65—
Results tracked5435

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama 4 Maverick: 26.6 (#324), Phi 3 Mini 4k Instruct: 26.6 (#323)

Coding benchmarks
BenchmarkLlama 4 MaverickPhi 3 Mini 4k Instruct
LMArena Coding13021093
SWE-bench Verified (bash only)21%—
Aider Polyglot15.6%—
SciCode33.1%—
WeirdML24.5%—
BigCodeBench Instruct49.7%—
LiveBench Coding—15.5%
BigCodeBench Complete61.4%—
ALE-Bench172.97—
HumanEval+—59.1%
MBPP+—54.2%

Agentic & Tool Use Not comparable

Llama 4 Maverick: 28.2 (#91), Phi 3 Mini 4k Instruct: —

Agentic & Tool Use benchmarks
BenchmarkLlama 4 MaverickPhi 3 Mini 4k Instruct
Berkeley Function Calling Leaderboard37.3%—

Reasoning Phi 3 Mini 4k Instruct leads

Llama 4 Maverick: 10.1 (#342), Phi 3 Mini 4k Instruct: 14.1 (#328)

Reasoning benchmarks
BenchmarkLlama 4 MaverickPhi 3 Mini 4k Instruct
LMArena Hard Prompts12811072
ARC-AGI-20%—
SimpleBench27.7%—
Kagi LLM Benchmark55.9%—
NYT Connections (extended)8%—
ARC-AGI-14.4%—
CritPt0%—
Chess Puzzles—0%
EnigmaEval0.6%—
LiveBench Reasoning—26.8%
DTBench61.9%—
LiveBench Data Analysis—34.7%
LMCA15.9%—
Adversarial NLI—52.8%
BIG-Bench Hard—71.7%
Epoch Capabilities Index132.2—
ForecastBench57.5—
HellaSwag—76.7%
LiveBench—22.4%
WinoGrande—70.8%

Math Too close to call

Llama 4 Maverick: 26.0 (#262), Phi 3 Mini 4k Instruct: 26.6 (#257)

Math benchmarks
BenchmarkLlama 4 MaverickPhi 3 Mini 4k Instruct
LMArena Math12991111
OTIS Mock AIME 2024-202520.6%—
Omni-MATH42.2%—
LiveBench Math—15.7%
MATH Level 573%—
FrontierMath (Feb 2025 set)0.7%—

Knowledge Llama 4 Maverick leads

Llama 4 Maverick: 33.4 (#204), Phi 3 Mini 4k Instruct: 28.5 (#246)

Knowledge benchmarks
BenchmarkLlama 4 MaverickPhi 3 Mini 4k Instruct
LMArena Expert12591045
GPQA Diamond67%—
Humanity's Last Exam5.7%—
MMLU-Pro81%—
Confabulations22.6%—
Vectara Hallucination Rate8.2%—
GPQA (HELM)65%—
ARC (AI2) Challenge—84.9%
MMLU—68.8%
OpenBookQA—88%
TriviaQA—64%

Multimodal Not comparable

Llama 4 Maverick: 31.6 (#105), Phi 3 Mini 4k Instruct: —

Multimodal benchmarks
BenchmarkLlama 4 MaverickPhi 3 Mini 4k Instruct
LMArena Vision1142—
GeoBench52%—
SpatialViz-Bench31.8%—

Multilingual Llama 4 Maverick leads

Llama 4 Maverick: 42.2 (#195), Phi 3 Mini 4k Instruct: 26.3 (#280)

Multilingual benchmarks
BenchmarkLlama 4 MaverickPhi 3 Mini 4k Instruct
LMArena Non-English12691021
LMArena Chinese12771021
LMArena French12591076
LMArena German12911044
LMArena Japanese1207935
LMArena Korean1203905
LMArena Russian12861022
LMArena Spanish12931085

Instruction Following Llama 4 Maverick leads

Llama 4 Maverick: 71.7 (#146), Phi 3 Mini 4k Instruct: 47.7 (#303)

Instruction Following benchmarks
BenchmarkLlama 4 MaverickPhi 3 Mini 4k Instruct
LMArena Instruction Following12671053
LiveBench Instruction Following—39.1%
IFEval90.8%—

Long Context Too close to call

Llama 4 Maverick: 31.4 (#279), Phi 3 Mini 4k Instruct: 31.7 (#276)

Long Context benchmarks
BenchmarkLlama 4 MaverickPhi 3 Mini 4k Instruct
LMArena Longer Query12801044
Fiction.LiveBench46.2%—

Writing & Preference Llama 4 Maverick leads

Llama 4 Maverick: 38.8 (#252), Phi 3 Mini 4k Instruct: 27.6 (#300)

Writing & Preference benchmarks
BenchmarkLlama 4 MaverickPhi 3 Mini 4k Instruct
LMArena Text12871073
LMArena Creative Writing12671037
LMArena Multi-Turn12891018
Short-Story Creative Writing62%—
EQ-Bench Creative Writing860—
WildBench80%—
LiveBench Language—9.2%

Frequently asked questions

Is Llama 4 Maverick better than Phi 3 Mini 4k Instruct?

Llama 4 Maverick is the stronger model overall, scoring 30.9 to 27.9 on the Noometry Index.

Is Llama 4 Maverick or Phi 3 Mini 4k Instruct better for coding?

They score almost the same on coding (26.6 vs 26.6); test both on your own repository before choosing.

How many benchmarks do Llama 4 Maverick and Phi 3 Mini 4k Instruct share?

17 benchmarks have published results for both models. Llama 4 Maverick has 54 scored results on Noometry and Phi 3 Mini 4k Instruct has 35.

Related comparisons

Go deeper