Guide · Local AI
How much VRAM do you need to run an LLM locally?
Add up three things: the weights (about 0.57 GB per billion parameters at the popular Q4_K_M quantisation), the KV cache that holds your context, and runtime overhead (we allow 10%, at least 1 GB). With an 8K context that comes to about 6.6 GB for an 8B model, 11 GB for 14B, 23 GB for 32B and 47 GB for 70B.
By Tahir NazirUpdated 12 min read
On this page
- How do you calculate the VRAM an LLM needs?
- How much VRAM does an 8B, 14B, 32B or 70B model need?
- Which quantisation should you use?
- How much VRAM does context length add?
- What fits on an 8, 12, 16, 24, 48 or 80 GB GPU?
- What if the model doesn’t fit on one GPU?
- Worked examples: RTX 4060, RTX 4090 and two RTX 3090s
- Questions people ask
How do you calculate the VRAM an LLM needs?
Add three numbers: the model’s weights, the KV cache for your context, and working space for the runtime. Our VRAM calculator uses this formula, and so does every figure here.
- Weights = parameters × bits per weight ÷ 8. Quantised files also store scales, so llama.cpp measures 4.89 bits per weight for Q4_K_M and 8.5 for Q8_0.
- KV cache = the keys and values every attention layer stores for each token in the context. It grows with context length and with parallel requests.
- Overhead = the GPU runtime, compute buffers and fragmentation. We assume 10% of the weights and cache combined, at least 1 GB. That’s an assumption; the real figure varies by runtime.
For example, Llama 3.1 8B at Q4_K_M with a 32K context needs 4.6 GB of weights, 4.0 GB of KV cache and 1.0 GB of overhead: 9.6 GB in total.
The weights are easy to check, because they’re the file you download. Our estimates land within 1.4% of Qwen’s official GGUF files:
| File | Estimate | Actual file |
|---|---|---|
Qwen3-8B-Q4_K_M.gguf | 4.66 GB | 4.68 GB |
Qwen3-14B-Q4_K_M.gguf | 8.41 GB | 8.38 GB |
Qwen3-32B-Q4_K_M.gguf | 18.65 GB | 18.40 GB |
Qwen3-32B-Q8_0.gguf | 32.42 GB | 32.43 GB |
File sizes from the Hugging Face API for the Qwen/Qwen3-8B, 14B and 32B GGUF repositories, checked 2026-10-08. GB here means 1,024³ bytes, the way GPU memory is quoted.
How much VRAM does an 8B, 14B, 32B or 70B model need?
Popular models at the three precisions people download, with a short and a long context. Each name links to that model’s full table.
| Model | Q4_K_M, 8K | Q4_K_M, 32K | Q8_0, 8K | Q8_0, 32K | FP16, 8K | FP16, 32K |
|---|---|---|---|---|---|---|
| Llama 3.1 8B | 6.6 GB | 9.6 GB | 9.9 GB | 13 GB | 18 GB | 21 GB |
| Qwen3 14B | 11 GB | 15 GB | 17 GB | 22 GB | 32 GB | 36 GB |
| Gemma 3 27B | 18 GB | 20 GB | 31 GB | 33 GB | 57 GB | 59 GB |
| Qwen3 30B-A3B | 20 GB | 22 GB | 34 GB | 37 GB | 63 GB | 66 GB |
| Qwen3 32B | 23 GB | 29 GB | 38 GB | 44 GB | 69 GB | 76 GB |
| Llama 3.3 70B | 47 GB | 55 GB | 80 GB | 88 GB | 147 GB | 156 GB |
Computed with our VRAM calculator’s formula from each model’s official config.json (fetched 2026-10-08).
Three patterns stand out. Precision matters most: Qwen3 32B needs about 3.1 times as much memory at FP16 as at Q4_K_M. Context is not free: going from 8K to 32K adds 6.6 GB to Qwen3 32B but only 2.1 GB to Gemma 3 27B, because 52 of Gemma’s 62 layers only look at the last 1,024 tokens. And mixture-of-experts models save compute, not memory: Qwen3 30B-A3B runs 8 of its 128 experts per token, but loads them all.
Which quantisation should you use?
Use Q4_K_M unless you have memory to spare; then step up to Q5_K_M, Q6_K or Q8_0. Storing each weight in fewer bits is the biggest lever you have: Qwen3 14B with a 32K context needs 36 GB at FP16 and 15 GB at Q4_K_M.
Qwen3 14B · 32K context · GB
- FP1616 bits/weight36 GB in total: 28 GB weights, 5.0 GB KV cache, 3.3 GB overhead
- Q8_08.5 bits/weight22 GB in total: 15 GB weights, 5.0 GB KV cache, 2.0 GB overhead
- Q6_K6.56 bits/weight18 GB in total: 11 GB weights, 5.0 GB KV cache, 1.6 GB overhead
- Q5_K_M5.7 bits/weight16 GB in total: 9.8 GB weights, 5.0 GB KV cache, 1.5 GB overhead
- Q4_K_M4.89 bits/weight15 GB in total: 8.4 GB weights, 5.0 GB KV cache, 1.3 GB overhead
- Q3_K_M4 bits/weight13 GB in total: 6.9 GB weights, 5.0 GB KV cache, 1.2 GB overhead
- Q2_K3.16 bits/weight11 GB in total: 5.4 GB weights, 5.0 GB KV cache, 1.0 GB overhead
llama.cpp’s maintainers measure the cost by comparing a quantised model’s predictions with the full-precision model’s on the same text. For Llama 3 8B:
| Format | Bits per weight | Perplexity increase | Same top token as FP16 |
|---|---|---|---|
| Q8_0 | 8.5 | +0.04% | 97.7% |
| Q6_K | 6.56 | +0.35% | 96.0% |
| Q4_K_M | 4.89 | +2.8% | 91.9% |
| Q2_K | 3.16 | +56.5% | 71.1% |
From the llama.cpp perplexity README (Wikitext-2, CUDA backend, no importance matrix). Perplexity measures how well a model predicts the next token, lower is better. “Same top token” is how often both models rank the same next token first.
Two caveats from the same source. The loss depends on the model: at Q4_K_M, Llama 3 8B’s perplexity rose 2.8% but Llama 2 7B’s only 1.4%. And perplexity is a proxy, not a task score, so test your own prompts. Files made with an importance matrix (imatrix) lose a little less.
- Q8_0: effectively the original model, at about half the FP16 size.
- Q6_K, Q5_K_M: a small step down, worth it when a few GB are spare.
- Q4_K_M: the usual default and best balance of size and quality.
- Q3_K_M, Q2_K: a last resort. At Q2_K, Llama 3 8B changed its top token almost three times in ten.
How much VRAM does context length add?
Every token of context adds a fixed amount to the KV cache, so a long context can need more memory than the model itself. For standard attention, each token costs 2 (a key and a value) × layers × KV heads × head size × 2 bytes at FP16.
| Model | Per 1,000 tokens | 8K | 32K | 128K |
|---|---|---|---|---|
| Llama 3.1 8B | 125 MB | 1.0 GB | 4.0 GB | 16 GB |
| Qwen3 32B | 250 MB | 2.0 GB | 8.0 GB | Over its native limit |
| Llama 3.3 70B | 313 MB | 2.5 GB | 10 GB | 40 GB |
| Gemma 3 27B | 78 MB | 1.0 GB | 2.9 GB | 10 GB |
| gpt-oss 20B | 23 MB | 0.2 GB | 0.8 GB | 3.0 GB |
Per 1,000 tokens counts only the layers that keep the whole context. Gemma 3 and gpt-oss use sliding-window layers, which stop growing once their window is full. Qwen3 handles 32K natively and reaches 128K only with YaRN scaling.
At its full 128K context, Llama 3.1 8B’s cache is 16 GB, 3.5 times its Q4_K_M weights. So set the context you need, not the maximum: measure a typical prompt with the token counter (see what a token is) and add room for the reply. Parallel requests multiply the cache: four 8K chats use as much as one 32K chat.
Quantise the KV cache
llama.cpp and Ollama can store the cache in 8 or 4 bits. An 8-bit (q8_0) cache takes 53% of the FP16 size and a 4-bit (q4_0) cache 28%. Ollama’s documentation says q8_0 “usually has no noticeable impact on the model’s quality”, while q4_0 has a loss “that may be more noticeable at higher context sizes”. Both need flash attention, which current versions turn on automatically where it’s supported.
# llama.cpp
llama-server -hf Qwen/Qwen3-14B-GGUF:Q4_K_M -c 32768 -ctk q8_0 -ctv q8_0
# Ollama (set when starting the server; applies to every model)
OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_CONTEXT_LENGTH=32768 ollama serveWhat fits on an 8, 12, 16, 24, 48 or 80 GB GPU?
“Fits” means the estimate is no bigger than the card’s memory, so leave some room for your desktop and other apps.
| VRAM | Example GPUs | Q4_K_M, 8K context | Q4_K_M, 32K context | Q8_0, 8K context |
|---|---|---|---|---|
| 8 GB | RTX 4060 | Qwen3 8B (6.8 GB) | Qwen2.5 7B (7.1 GB) | Gemma 3 4B (5.5 GB) |
| 12 GB | RTX 3060, RTX 4070, RTX 5070 | Phi-4 14B (11 GB) | Gemma 3 12B (10 GB) | Qwen3 8B (10 GB) |
| 16 GB | RTX 4060 Ti 16GB, RTX 4080 Super, RTX 5060 Ti 16GB | Phi-4 14B (11 GB) | Qwen3 14B (15 GB) | Mistral Nemo 12B (15 GB) |
| 24 GB | RTX 3090, RTX 4090, Radeon RX 7900 XTX | Qwen3 32B (23 GB) | Gemma 4 31B (23 GB) | Phi-4 14B (18 GB) |
| 48 GB | L40S | Llama 3.3 70B (47 GB) | Qwen3 32B (29 GB) | Qwen3 32B (38 GB) |
| 80 GB | A100 80GB, H100 80GB | Llama 4 Scout (70 GB) | Llama 4 Scout (71 GB) | Llama 3.3 70B (80 GB) |
Same estimates as the VRAM calculator, FP16 KV cache, one request. Models whose maximum context is under 32K are left out of that column, and gpt-oss is covered below. GPU memory sizes from NVIDIA and AMD product pages.
The big steps are 8 to 12 GB, which takes you from 8B to 14B models, and 16 to 24 GB, which opens up the 27B to 32B class. A 70B model at Q4_K_M (47 GB) needs a 48 GB card or two 24 GB cards. At FP16 it needs 147 GB, more than an 80 GB card.
OpenAI’s gpt-oss is the exception to these quantisation levels. It is released in MXFP4, a 4-bit format, and its GGUF files are about the same size whichever label they carry, so go by the official files: 11.3 GB for gpt-oss 20B, about 13 GB in total with a 32K context, which fits a 16 GB card; and 59 GB for gpt-oss 120B, about 66 GB, which fits an 80 GB card.
What if the model doesn’t fit on one GPU?
You can split it across several GPUs or keep part of it in system RAM. Splitting across GPUs keeps everything in fast VRAM, so it’s the better option when you can.
Several GPUs
llama.cpp uses every visible GPU by default. Its default split mode, layer, puts whole layers and their share of the KV cache on each card and runs them in turn, so for a single request two GPUs add memory rather than speed. --tensor-split 3,1 sets the proportions when the cards differ. Its row and experimental tensor modes, like vLLM’s tensor parallelism (--tensor-parallel-size 2), instead split each layer across GPUs so they work at the same time. Memory adds up: Llama 3.3 70B at Q4_K_M needs 2 × 24 GB cards, and each card keeps its own buffers, so leave a little extra headroom.
CPU offload
llama.cpp and Ollama can keep some layers in system RAM and run them on the CPU. In llama.cpp, -ngl (--n-gpu-layers) sets how many layers go to the GPU; its default, auto, picks what fits your VRAM. Ollama also decides for you, and ollama ps shows the split, for example “48%/52% CPU/GPU”. Every token passes through every layer, and system RAM is much slower than VRAM, so offloaded layers slow generation down. Try a smaller quantisation or a shorter context first.
Mixture-of-experts models offload better, because only a few experts run for each token. llama.cpp’s --n-cpu-moe N keeps the expert weights of the first N layers in system RAM and leaves attention on the GPU. That suits models like gpt-oss 120B, which runs 4 of its 128 experts per token but whose official file alone is 59 GB.
Worked examples: RTX 4060, RTX 4090 and two RTX 3090s
“I have an RTX 4060 with 8 GB”
Run 7B to 8B models at Q4_K_M. Llama 3.1 8B needs 6.6 GB and Qwen3 8B 6.8 GB with an 8K context. At 32K, Llama 3.1 8B needs 9.6 GB, but an 8-bit KV cache brings it to 7.7 GB, which just fits. Gemma 3 12B at Q4_K_M (8.8 GB) is already over the limit, so 12B to 14B models need partial CPU offload or a 12 GB card.
“I have an RTX 4090 with 24 GB”
24 GB suits 24B to 32B models at Q4_K_M. Qwen3 32B needs 22.7 GB at 8K, a tight fit. At 32K it needs 29 GB, still 25 GB with an 8-bit cache, so 16K with an 8-bit cache (22.9 GB) is its practical limit. For long contexts, Gemma 3 27B (20 GB at 32K) and Qwen3 30B-A3B (22 GB) fit. A 70B model (47 GB) doesn’t.
“I have two RTX 3090s (48 GB in total)”
This is a common home setup for 70B models. Llama 3.3 70B at Q4_K_M needs 46.9 GB of the 48 GB with an 8K context, so there is little room to spare. A 16K window with an 8-bit cache needs 47.1 GB; for 32K, drop to Q3_K_M (42.0 GB with an 8-bit cache) and accept some quality loss. Or run Qwen3 32B at near-lossless Q8_0 (38 GB).
Free toolLLM VRAM calculatorPick a model, quantisation, context length and KV cache type to see the breakdown and which GPUs and Macs can run it.Macs work differently, because the GPU shares system memory: see how to run LLMs locally on a Mac. To compare with a hosted API, use the LLM cost calculator.
FAQ
Questions people ask
How much VRAM do I need to run a 70B model?
About 47 GB for Llama 3.3 70B at Q4_K_M with an 8K context, 80 GB at Q8_0 and 147 GB at FP16. At Q4_K_M that means two 24 GB cards or one 48 GB card. A 32K context raises the Q4_K_M figure to 55 GB, so long contexts need an 8-bit KV cache or a smaller quantisation.
Can I run an LLM on 8 GB of VRAM?
Yes. 7B and 8B models at Q4_K_M fit with an 8K context: Llama 3.1 8B needs about 6.6 GB. For a 32K context, quantise the KV cache to 8-bit (about 7.7 GB in total). Models from 12B up need a smaller quantisation or partial CPU offload, which is slower.
Is Q4 good enough, or should I use Q8?
For most uses Q4_K_M is good enough, which is why it’s the usual default. In llama.cpp’s measurements on Llama 3 8B, Q4_K_M raised perplexity by 2.8% and Q8_0 by 0.04%. If Q8_0 fits with the context you need, use it; if it doesn’t, Q5_K_M or Q6_K are good middle steps.
Why does my model use more VRAM than its file size?
The file holds only the weights. At run time the runtime also allocates the KV cache for the whole context window up front, plus compute buffers. A long default context, or several parallel requests, can add gigabytes. Lower the context length or quantise the cache to bring it down.
Can I use system RAM instead of VRAM?
Partly. llama.cpp and Ollama can run some layers from system RAM on the CPU, so a model slightly too big for your card still loads. Generation slows down because system RAM is much slower than VRAM. Mixture-of-experts models cope best, since only a few experts run for each token.
Do mixture-of-experts models need less VRAM?
No. Every expert must be in memory, so Qwen3 30B-A3B needs about as much as a dense 30B model (20 GB at Q4_K_M with 8K). What they save is compute: only 8 of its 128 experts run for each token, so it generates text much faster than a dense model of the same size.
Try it
Tools from this guide
Keep reading