Model comparison

Llama2 70b Steerlm Chat vs Phi-4 Mini

Llama2 70b Steerlm Chat and Phi-4 Mini score almost the same on the Noometry Index (31.8 vs 30.9), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Side by side

Llama2 70b Steerlm Chat and Phi-4 Mini specifications
Llama2 70b Steerlm ChatPhi-4 Mini
ProviderNVIDIAMicrosoft
Noometry Index31.830.9
Released—2024-12-11
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.075
Output $ / M tokens—$0.30
Results tracked93

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 29.9 (#300), Phi-4 Mini: 28.1 (#317)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatPhi-4 Mini
SciCode—10.8%
LMArena Coding1025—

Reasoning Phi-4 Mini leads

Llama2 70b Steerlm Chat: 20.0 (#246), Phi-4 Mini: 22.4 (#195)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatPhi-4 Mini
CritPt—0%
LMArena Hard Prompts1047—

Math Not comparable

Llama2 70b Steerlm Chat: 31.3 (#226), Phi-4 Mini: —

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatPhi-4 Mini
LMArena Math1072—

Knowledge Not comparable

Llama2 70b Steerlm Chat: —, Phi-4 Mini: 25.3 (#262)

Knowledge benchmarks
BenchmarkLlama2 70b Steerlm ChatPhi-4 Mini
Vectara Hallucination Rate—23.5%

Multilingual Not comparable

Llama2 70b Steerlm Chat: 28.8 (#270), Phi-4 Mini: —

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatPhi-4 Mini
LMArena Non-English1063—

Instruction Following Not comparable

Llama2 70b Steerlm Chat: 54.2 (#279), Phi-4 Mini: —

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatPhi-4 Mini
LMArena Instruction Following1060—

Long Context Not comparable

Llama2 70b Steerlm Chat: 30.4 (#288), Phi-4 Mini: —

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatPhi-4 Mini
LMArena Longer Query998—

Writing & Preference Not comparable

Llama2 70b Steerlm Chat: 31.6 (#283), Phi-4 Mini: —

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatPhi-4 Mini
LMArena Text1098—
LMArena Creative Writing1091—
LMArena Multi-Turn1058—

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Phi-4 Mini?

Llama2 70b Steerlm Chat and Phi-4 Mini score almost the same on the Noometry Index (31.8 vs 30.9), so choose on price, context window or the category you care about most.

Is Llama2 70b Steerlm Chat or Phi-4 Mini better for coding?

Llama2 70b Steerlm Chat scores higher on coding benchmarks: 29.9 versus 28.1 in the Noometry coding category.

How many benchmarks do Llama2 70b Steerlm Chat and Phi-4 Mini share?

0 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Phi-4 Mini has 3.

Related comparisons

Go deeper