Tokens & Costs
LLM VRAM calculator: can your GPU run it?
Work out how much GPU memory a local model needs (weights, KV cache and overhead) and which graphics cards or Macs can run it. Free, and it runs in your browser.
8.0B parameters · 131,072-token context · model card
GGUF; the most popular balance of size and quality.
Requests served at the same time.
- Weights
- 4.6 GB
- KV cache
- 1.0 GB
- Overhead
- 1.0 GB
| GPU | Memory | Runs it? |
|---|---|---|
| GeForce RTX 4060 | 8 GB | Yes |
| GeForce RTX 3060 | 12 GB | Yes |
| GeForce RTX 4070 | 12 GB | Yes |
| GeForce RTX 5070 | 12 GB | Yes |
| GeForce RTX 4060 Ti 16GB | 16 GB | Yes |
| GeForce RTX 4080 Super | 16 GB | Yes |
| GeForce RTX 5060 Ti 16GB | 16 GB | Yes |
| GeForce RTX 5070 Ti | 16 GB | Yes |
| GeForce RTX 5080 | 16 GB | Yes |
| GeForce RTX 3090 | 24 GB | Yes |
| GeForce RTX 4090 | 24 GB | Yes |
| Radeon RX 7900 XTX | 24 GB | Yes |
| GeForce RTX 5090 | 32 GB | Yes |
| L4 | 24 GB | Yes |
| L40S | 48 GB | Yes |
| A100 80GB | 80 GB | Yes |
| H100 80GB | 80 GB | Yes |
| RTX PRO 6000 Blackwell | 96 GB | Yes |
| H200 | 141 GB | Yes |
| Instinct MI300X | 192 GB | Yes |
Apple Silicon Macs (unified memory)
- 16 GB: fits
- 24 GB: fits
- 32 GB: fits
- 36 GB: fits
- 48 GB: fits
- 64 GB: fits
- 96 GB: fits
- 128 GB: fits
- 192 GB: fits
- 256 GB: fits
- 512 GB: fits
Filled: fits as is. Outlined: fits after raising the GPU memory limit (leaving 8 GB for macOS). macOS gives the GPU two-thirds of memory up to 32 GB and three-quarters above by default.
Steps
How to use the LLM VRAM calculator
- Pick a model, or choose “Custom size” and enter its parameter count.
- Choose the quantisation you plan to download, such as Q4_K_M.
- Set the context length you need, and the batch size if you serve several requests at once.
- Read the total and its breakdown, then check which GPUs and Mac memory sizes can run it.
Method
How it works
Running a model locally needs memory for three things: the model’s weights, the KV cache that holds the conversation, and working space for the runtime. The calculator adds them up the same way the runtimes allocate them.
The formula
total = weights + KV cache + overhead
- Weights = parameters × bits per weight ÷ 8. The bits per weight are llama.cpp’s measured figures, which include the scales that quantised formats store alongside the weights: for example 4.89 for Q4_K_M and 8.5 for Q8_0. They were measured on Llama 3.1 8B; mixed formats vary slightly between models.
- KV cache = for each attention layer, the values it stores per token × the tokens it keeps × the batch size × bytes per value. A standard layer stores keys and values for each of its KV heads (2 × KV heads × head size). DeepSeek-style MLA layers store one compressed vector instead. Sliding-window and chunked layers only keep their last few thousand tokens, and Mamba or linear-attention layers keep a small fixed state rather than a cache.
- Overhead = 10% of weights and cache, at least 1 GB, for the CUDA or Metal context, compute buffers and fragmentation. This is an assumption, not a measurement: the real figure depends on the runtime and its settings.
Where the numbers come from
Layer counts, KV heads, head sizes, attention windows and expert counts are read from each model’s official config.json on Hugging Face, pinned to a specific version, and parameter counts come from the published weights. Architectures the calculator can’t read reliably, such as the newest sparse-attention designs, are left out rather than estimated. GB here means 1,024³ bytes, the way GPU memory is quoted.
Reading the result
A model “fits” a GPU when the estimate is no larger than its memory. In practice, leave a little room: your desktop and other apps also use GPU memory, and long prompts need larger compute buffers. If you’re close to the limit, use a smaller quantisation, quantise the KV cache to 8-bit, or shorten the context. For a cost comparison with hosted models, see the LLM API cost calculator.
Examples
Worked examples
Memory needed at an 8K context, batch 1
| Model | FP16 / BF16 | Q8_0 (8-bit) | Q4_K_M |
|---|---|---|---|
| Llama 3.1 8B | 18 GB | 9.9 GB | 6.6 GB |
| Qwen3 14B | 32 GB | 17 GB | 11 GB |
| Gemma 3 27B | 57 GB | 31 GB | 18 GB |
| Qwen3 32B | 69 GB | 38 GB | 23 GB |
| gpt-oss 20B | 13 GB | 13 GB | 13 GB |
| Qwen3 30B-A3B | 63 GB | 34 GB | 20 GB |
| Llama 3.3 70B | 147 GB | 80 GB | 47 GB |
| gpt-oss 120B | 65 GB | 65 GB | 65 GB |
| Qwen3 235B-A22B | 483 GB | 258 GB | 149 GB |
| DeepSeek R1 (671B) | 1375 GB | 731 GB | 421 GB |
Weights + FP16 KV cache + overhead, from the formula above. Model configs fetched 2026-10-08.
FAQ
Frequently asked questions
How much VRAM do I need for a 70B model?
For Llama 3.3 70B with an 8K context, about 47 GB at Q4_K_M, 80 GB at 8-bit and 147 GB at full 16-bit precision. At Q4 that needs 2 × 24 GB cards or one 48 GB card, or a Mac with 64 GB of memory. Longer contexts add more: at 128K the cache alone is 40 GB.
Can I run an 8B model on an 8 GB GPU?
Yes, with quantisation. Llama 3.1 8B at Q4_K_M with an 8K context needs about 6.6 GB, which fits an 8 GB card. At full precision it needs about 18 GB, so a 24 GB card.
What is quantisation, and how much quality does it cost?
Quantisation stores each weight in fewer bits, for example about 4.9 bits instead of 16, which cuts memory by roughly two-thirds. 8-bit is close to lossless, Q5 and Q6 lose very little, and Q4_K_M is the usual sweet spot. Below 4 bits quality drops more quickly, especially for smaller models.
Why does a longer context need more memory?
For every token in the context, each attention layer keeps its keys and values in the KV cache, so the cache grows in proportion to the context length and the batch size. Models that use grouped-query attention, sliding windows or compressed (MLA) attention keep much smaller caches, which this calculator reads from each model’s configuration. Quantising the cache to 8-bit roughly halves it.
Can I split a model across two GPUs?
Yes. llama.cpp, vLLM and similar runtimes can split the layers across several GPUs, so two 24 GB cards hold about 48 GB of model. Each GPU needs a little extra for its own buffers, and splitting can be slower than one big card, but it is the usual way to run 70B models at home.
What if the model doesn’t fit in VRAM?
Runtimes such as llama.cpp and Ollama can keep some layers in normal system RAM and run them on the CPU. It works, but every layer on the CPU slows generation down considerably. A smaller quantisation, a shorter context or a smaller model usually gives a better experience.
How much of a Mac’s memory can the GPU use?
By default macOS lets the GPU use about two-thirds of unified memory on Macs with up to 32 GB and three-quarters on larger ones. You can raise the limit with the iogpu.wired_limit_mb setting; the calculator assumes you keep 8 GB for macOS when you do.
Do mixture-of-experts models need less memory?
No: every expert has to be loaded, so a 30B mixture-of-experts model needs as much memory as a 30B dense model. What it saves is compute: only a few experts run for each token, so it generates text much faster than a dense model of the same size.
Related
Related tools
- AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- AI Model ComparisonCompare prices, context windows and features across models.
- AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.
- Context Window CheckerSee whether your text fits each model's context window.
- Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.