AI glossary · Local AI
What is VRAM, and why does it limit local LLMs?
Also called: video RAM, GPU memory, graphics memory
Definition
VRAM is the memory on a graphics card, and for running AI models locally it is the main limit: a model runs at full GPU speed only when its weights, KV cache and working buffers all fit in it.
Explained
How it works
VRAM (video RAM) sits on the graphics card itself, separate from your computer’s system RAM. A model runs fastest when everything it needs is there. When it doesn’t fit, llama.cpp and Ollama can keep some layers in system RAM and run them on the CPU, which works but is much slower.
Three things fill it: the weights (parameters × bits per weight ÷ 8), the KV cache for your context, and working space for the runtime. Our calculator estimates that last part as 10% of the other two, at least 1 GB, which is an assumption rather than a measurement.
Apple Silicon Macs have no separate VRAM: the CPU and GPU share unified memory. By default macOS lets the GPU use only part of it: Apple’s figures for M1 Pro and M1 Max MacBook Pros are 21 GB of 32 GB and 48 GB of 64 GB, and llama.cpp users report about two-thirds on Macs with up to 32 GB and three-quarters above. You can raise the limit with sudo sysctl iogpu.wired_limit_mb=…; our calculator then keeps 8 GB free for macOS.
Example
Qwen3 32B at Q4_K_M on one GPU
Qwen3 32B at Q4_K_M has 19 GB of weights. With an 8K context and a batch of one, the cache and overhead bring it to 23 GB, which fits a 24 GB card such as the GeForce RTX 3090. At 32K the cache alone is 8.0 GB and the 29 GB total needs a 32 GB card. A 16K context with an 8-bit cache fits 24 GB again. On a Mac, the 8K setting needs 36 GB of unified memory with the default limit.
| Setting | Weights | KV cache | Overhead | Total | Smallest single GPU |
|---|---|---|---|---|---|
| 8K context, FP16 cache | 18.7 GB | 2.0 GB | 2.1 GB | 22.7 GB | 24 GB (RTX 3090) |
| 32K context, FP16 cache | 18.7 GB | 8.0 GB | 2.7 GB | 29.3 GB | 32 GB (RTX 5090) |
| 16K context, q8_0 cache | 18.7 GB | 2.1 GB | 2.1 GB | 22.9 GB | 24 GB (RTX 3090) |
From our VRAM calculator, batch 1. GPUs from the list it checks; leave some headroom for your desktop and other apps. Model config fetched 2026-10-08. GB = 1,024³ bytes, the way GPU memory is quoted.
Cost and quality
Why it matters
Memory, more than raw GPU speed, decides which models you can run locally: too little and the model either won’t load or spills onto the much slower CPU. That is why quantisation and context length matter so much: both are ways of fitting the same model into less memory.
When buying hardware, the memory size sets the largest model you can run. Check a specific model and context in our VRAM calculator before you buy.
Don’t mix up
Common confusions
- VRAM vs system RAM
- System RAM is the memory on your motherboard. A PC with 64 GB of RAM and an 8 GB graphics card can hold at most 8 GB of weights, cache and buffers on the GPU; any layers that don’t fit run on the CPU, which slows the whole model down.
- Unified memory vs VRAM
- A Mac’s memory is shared, so a 32 GB Mac doesn’t give the GPU 32 GB. With the default limit it gets about 21 GB, and macOS and your apps need the rest.
Go deeper
Try it and read more
- Free toolGPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.
- Free toolQwen3 32B VRAM requirementsQwen3 32B needs about 23 GB of VRAM at Q4_K_M and 38 GB at 8-bit with an 8K context.
- Guide · 12 min readHow much VRAM do you need to run an LLM locally?VRAM needed for 8B to 70B models at Q4, Q8 and FP16, what fits on 8 to 80 GB GPUs, and how context length, KV cache and CPU offload change it.
- Guide · 12 min readHow to run LLMs locally on a MacWhat fits on 16 to 128 GB Apple Silicon Macs, how much memory the GPU can use, how to set up Ollama, LM Studio, llama.cpp or MLX, and how fast it runs.
Related
Related terms
- KV cacheThe KV cache is the memory where a language model keeps the attention keys and values of every token it has already processed, so each new token is computed without reprocessing the whole sequence.
- QuantisationQuantisation is storing a model’s weights in fewer bits than the 16 per weight most models are released with, so the model needs less memory and runs on smaller hardware, at a small cost in accuracy.
- Mixture of expertsA mixture-of-experts (MoE) model splits parts of each layer into many parallel sub-networks called experts and uses a small router to send each token through only a few of them, so it computes with a fraction of its parameters.
- GGUFGGUF is a single-file binary format from the ggml and llama.cpp project that stores a model’s weights together with everything needed to run them, such as its architecture, tokenizer and quantisation type.
- Open weightsOpen weights means a model’s trained parameters are published for anyone to download and run on their own hardware, under a licence that sets what they may do with them.