AI glossary · Local AI
What is the KV cache in an LLM?
Also called: key-value cache, attention cache
Definition
The KV cache is the memory where a language model keeps the attention keys and values of every token it has already processed, so each new token is computed without reprocessing the whole sequence.
Explained
How it works
Each attention layer turns every token into a query, a key and a value. To write the next token, the model compares its new query with the keys of all earlier tokens and mixes their values. Those keys and values don’t change, so runtimes keep them in memory instead of recomputing them at every step.
It saves computation, not memory: every new token still reads the whole cache, and the cache grows with every token. Its size is 2 (keys and values) × layers × KV heads × head size × bytes per value × tokens, multiplied by the batch size, because each sequence has its own. Grouped-query attention, sliding-window layers and compressed attention all shrink it, so the real figure depends on each model’s config.
Runtimes can store it at lower precision. llama.cpp has --cache-type-k and --cache-type-v, and Ollama has OLLAMA_KV_CACHE_TYPE; both accept q8_0 and q4_0. A quantised cache needs Flash Attention: llama.cpp refuses a quantised V cache without it, and Ollama applies cache quantisation only when Flash Attention is on, which it enables automatically where the hardware supports it. Ollama’s docs put q8_0 at about half the memory of f16, with a very small loss of precision.
Example
Llama 3.1 8B at 8K and 128K tokens
Llama 3.1 8B has 32 layers, each with 8 KV heads of 128 values. In FP16 every token adds 2 × 32 × 8 × 128 × 2 bytes = 128 KiB. At 8K tokens the cache is 1.0 GB. At the model’s full 128K context it is 16 GB, about 3.5 times the 4.6 GB of Q4_K_M weights.
| Context | KV cache (FP16) | KV cache (q8_0) | Total at Q4_K_M, FP16 cache |
|---|---|---|---|
| 8,192 tokens | 1.0 GB | 0.5 GB | 6.6 GB |
| 32,768 tokens | 4.0 GB | 2.1 GB | 9.6 GB |
| 131,072 tokens | 16.0 GB | 8.5 GB | 22.6 GB |
From our VRAM calculator: batch 1, Q4_K_M weights, overhead 10% (at least 1 GB). q8_0 stores 32 values in 34 bytes. Architecture from the model’s config.json, fetched 2026-10-08. GB = 1,024³ bytes.
Cost and quality
Why it matters
On your own GPU, the KV cache is often what turns a model that fits into one that doesn’t. The weights are fixed, but the cache grows with the context length and with every parallel request, so a long context can need more memory than the model itself.
When memory is tight, set a shorter context than the maximum, quantise the cache to q8_0, or pick a model with fewer KV heads or sliding-window layers. Our VRAM calculator shows the cache separately for each model.
Don’t mix up
Common confusions
- KV cache vs prompt caching
- The KV cache is the working memory of whatever hardware runs the model; local servers such as llama.cpp’s can reuse it for a shared prefix in the next request. Prompt caching is the API version of that idea: the provider keeps a processed prefix between requests and bills it at a discount.
- KV cache vs context window
- The context window is the most tokens a model can handle. The KV cache is the memory the tokens you actually use take up, which is why local runtimes let you set a shorter context to save memory.
Go deeper
Try it and read more
- Free toolGPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.
- Free toolContext Window CheckerSee whether your text fits each model’s context window.
- Guide · 12 min readHow much VRAM do you need to run an LLM locally?VRAM needed for 8B to 70B models at Q4, Q8 and FP16, what fits on 8 to 80 GB GPUs, and how context length, KV cache and CPU offload change it.
Related
Related terms
- Prompt cachingPrompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.
- Context windowA context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.
- VRAMVRAM is the memory on a graphics card, and for running AI models locally it is the main limit: a model runs at full GPU speed only when its weights, KV cache and working buffers all fit in it.
- QuantisationQuantisation is storing a model’s weights in fewer bits than the 16 per weight most models are released with, so the model needs less memory and runs on smaller hardware, at a small cost in accuracy.
- Time to first tokenTime to first token (TTFT) is how long a model takes from receiving a request to returning the first token of its reply, the wait a user feels before streamed text starts to appear.