Model comparison

Mistral 7B vs Step 3.7 Flash

Step 3.7 Flash is the stronger model overall, scoring 37.3 to 23.0 on the Noometry Index. Mistral 7B costs 1.7× less per token, which makes it the better buy when Step 3.7 Flash's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Step 3.7 Flash StepFun

37.3

Rank #207 Reported

Summary

  • The widest gap is in math, where Step 3.7 Flash leads 42.9 to 8.1.
  • Mistral 7B is cheaper at $0.25 / $0.25 per million input/output tokens, against $0.18 / $1.11 for Step 3.7 Flash.
  • Step 3.7 Flash accepts more context: 256K tokens versus 8K.

Side by side

Mistral 7B and Step 3.7 Flash specifications
Mistral 7BStep 3.7 Flash
ProviderMistral AIStepFun
Noometry Index23.037.3
Released2023-09-272026-05-29
WeightsOpenOpen
Context window8K256K
Max output8K256K
Input $ / M tokens$0.25$0.18
Output $ / M tokens$0.25$1.11
Results tracked375

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 3.7 Flash leads

Mistral 7B: 26.4 (#326), Step 3.7 Flash: 40.0 (#150)

Coding benchmarks
BenchmarkMistral 7BStep 3.7 Flash
SciCode—40%
BigCodeBench Instruct19.5%—
LMArena Coding1082—
BigCodeBench Complete27.3%—
ALE-Bench—694.12
HumanEval+36%—
MBPP+42.1%—

Reasoning Step 3.7 Flash leads

Mistral 7B: 13.1 (#336), Step 3.7 Flash: 21.6 (#219)

Reasoning benchmarks
BenchmarkMistral 7BStep 3.7 Flash
NYT Connections (extended)—39.7%
CritPt—2.3%
Chess Puzzles0%—
LMArena Hard Prompts1067—
DTBench42.5%—
Adversarial NLI47.1%—
BIG-Bench Hard56.1%—
Epoch Capabilities Index112.21—
HellaSwag81%—
PIQA83%—
WinoGrande75.3%—

Math Step 3.7 Flash leads

Mistral 7B: 8.1 (#325), Step 3.7 Flash: 42.9 (#82)

Math benchmarks
BenchmarkMistral 7BStep 3.7 Flash
MathArena Final-Answer Competitions—68.5%
OTIS Mock AIME 2024-20250.3%—
LMArena Math1085—
MATH Level 53.7%—
GSM8K54.4%—

Knowledge Not comparable

Mistral 7B: 7.4 (#311), Step 3.7 Flash: —

Knowledge benchmarks
BenchmarkMistral 7BStep 3.7 Flash
GPQA Diamond15.2%—
LMArena Expert1036—
ARC (AI2) Challenge78.6%—
BoolQ87.4%—
MMLU62.5%—
OpenBookQA79.8%—
TriviaQA75.2%—

Multilingual Not comparable

Mistral 7B: 25.8 (#283), Step 3.7 Flash: —

Multilingual benchmarks
BenchmarkMistral 7BStep 3.7 Flash
LMArena Non-English1012—
LMArena Chinese1009—
LMArena French1037—
LMArena German987—
LMArena Japanese878—
LMArena Russian1018—
LMArena Spanish1026—

Instruction Following Not comparable

Mistral 7B: 54.2 (#280), Step 3.7 Flash: —

Instruction Following benchmarks
BenchmarkMistral 7BStep 3.7 Flash
LMArena Instruction Following1060—

Long Context Not comparable

Mistral 7B: 32.2 (#271), Step 3.7 Flash: —

Long Context benchmarks
BenchmarkMistral 7BStep 3.7 Flash
LMArena Longer Query1060—

Writing & Preference Not comparable

Mistral 7B: 30.7 (#286), Step 3.7 Flash: —

Writing & Preference benchmarks
BenchmarkMistral 7BStep 3.7 Flash
LMArena Text1090—
LMArena Creative Writing1068—
LMArena Multi-Turn1062—

Frequently asked questions

Is Mistral 7B better than Step 3.7 Flash?

Step 3.7 Flash is the stronger model overall, scoring 37.3 to 23.0 on the Noometry Index. Mistral 7B costs 1.7× less per token, which makes it the better buy when Step 3.7 Flash's lead doesn't matter for your workload.

Which is cheaper, Mistral 7B or Step 3.7 Flash?

Mistral 7B is cheaper. It lists at $0.25 per million input tokens and $0.25 per million output tokens; Step 3.7 Flash lists at $0.18 and $1.11.

Is Mistral 7B or Step 3.7 Flash better for coding?

Step 3.7 Flash scores higher on coding benchmarks: 40.0 versus 26.4 in the Noometry coding category.

Which has the bigger context window?

Step 3.7 Flash does, with 256K tokens against 8K.

How many benchmarks do Mistral 7B and Step 3.7 Flash share?

0 benchmarks have published results for both models. Mistral 7B has 37 scored results on Noometry and Step 3.7 Flash has 5.

Related comparisons

Go deeper