Free tool
LLM VRAM Calculator: Can I Run It Locally?
Enter a model's size, quantization and context length. The estimate adds the weights, the KV cache for your context and runtime overhead, then lists hardware with enough memory. Treat it as a planning figure: real use varies by ±15% with the runtime.
≈ 6.9 GB (weights 4.8 + KV cache 0.9 + overhead 1.2)
Fits on: RTX 4060 Ti 16 GB, RTX 3090 / 4090 (24 GB), RTX 5090 (32 GB), Mac with 64 GB unified memory
Looking for a model to run? See the best open-weight models and the best local model for coding.
Frequently asked questions
How much VRAM do I need to run a 70B model?
About 49 GB at 4-bit with an 8K context, or 155 GB at 16-bit.
How much VRAM does an 8B model need?
About 6.9 GB at 4-bit with an 8K context, which fits a 12–16 GB card.
What about mixture-of-experts models?
Enter the total parameter count: every expert has to be in memory even though only some run per token. Offloading experts to system RAM lowers VRAM needs at a speed cost.