Free tool

# LLM VRAM Calculator: Can I Run It Locally?

> Estimate GPU memory needed to run an open-weight LLM locally by parameter count, quantization and context length, and see which GPUs fit.
- Canonical page: https://noometry.com/tools/vram-calculator
- Title: LLM VRAM Calculator: Can I Run It Locally? | Noometry

Enter a model's size, quantization and context length. The estimate adds the weights, the KV cache for your context and runtime overhead, then lists hardware with enough memory. Treat it as a planning figure: real use varies by ±15% with the runtime.

Parameters (billions)

Quantization 

Context length (tokens) 

 8-bit KV cache

≈ **6.9 GB** (weights 4.8 + KV cache 0.9 + overhead 1.2)

Fits on: RTX 4060 Ti 16 GB, RTX 3090 / 4090 (24 GB), RTX 5090 (32 GB), Mac with 64 GB unified memory

Looking for a model to run? See the [best open-weight models](https://noometry.com/best/open-source) and the [best local model for coding](https://noometry.com/best/for/local-code).

## Frequently asked questions

### How much VRAM do I need to run a 70B model?

About 49 GB at 4-bit with an 8K context, or 155 GB at 16-bit.

### How much VRAM does an 8B model need?

About 6.9 GB at 4-bit with an 8K context, which fits a 12–16 GB card.

### What about mixture-of-experts models?

Enter the total parameter count: every expert has to be in memory even though only some run per token. Offloading experts to system RAM lowers VRAM needs at a speed cost.
