Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Local AI

What is the KV cache in an LLM?

Also called: key-value cache, attention cache

Definition

The KV cache is the memory where a language model keeps the attention keys and values of every token it has already processed, so each new token is computed without reprocessing the whole sequence.

Explained

How it works

Each attention layer turns every token into a query, a key and a value. To write the next token, the model compares its new query with the keys of all earlier tokens and mixes their values. Those keys and values don’t change, so runtimes keep them in memory instead of recomputing them at every step.

It saves computation, not memory: every new token still reads the whole cache, and the cache grows with every token. Its size is 2 (keys and values) × layers × KV heads × head size × bytes per value × tokens, multiplied by the batch size, because each sequence has its own. Grouped-query attention, sliding-window layers and compressed attention all shrink it, so the real figure depends on each model’s config.

Runtimes can store it at lower precision. llama.cpp has --cache-type-k and --cache-type-v, and Ollama has OLLAMA_KV_CACHE_TYPE; both accept q8_0 and q4_0. A quantised cache needs Flash Attention: llama.cpp refuses a quantised V cache without it, and Ollama applies cache quantisation only when Flash Attention is on, which it enables automatically where the hardware supports it. Ollama’s docs put q8_0 at about half the memory of f16, with a very small loss of precision.

Example

Llama 3.1 8B at 8K and 128K tokens

Llama 3.1 8B has 32 layers, each with 8 KV heads of 128 values. In FP16 every token adds 2 × 32 × 8 × 128 × 2 bytes = 128 KiB. At 8K tokens the cache is 1.0 GB. At the model’s full 128K context it is 16 GB, about 3.5 times the 4.6 GB of Q4_K_M weights.

Llama 3.1 8B: KV cache by context length and cache type
ContextKV cache (FP16)KV cache (q8_0)Total at Q4_K_M, FP16 cache
8,192 tokens1.0 GB0.5 GB6.6 GB
32,768 tokens4.0 GB2.1 GB9.6 GB
131,072 tokens16.0 GB8.5 GB22.6 GB

From our VRAM calculator: batch 1, Q4_K_M weights, overhead 10% (at least 1 GB). q8_0 stores 32 values in 34 bytes. Architecture from the model’s config.json, fetched 2026-10-08. GB = 1,024³ bytes.

Cost and quality

Why it matters

On your own GPU, the KV cache is often what turns a model that fits into one that doesn’t. The weights are fixed, but the cache grows with the context length and with every parallel request, so a long context can need more memory than the model itself.

When memory is tight, set a shorter context than the maximum, quantise the cache to q8_0, or pick a model with fewer KV heads or sliding-window layers. Our VRAM calculator shows the cache separately for each model.

Don’t mix up

Common confusions

KV cache vs prompt caching
The KV cache is the working memory of whatever hardware runs the model; local servers such as llama.cpp’s can reuse it for a shared prefix in the next request. Prompt caching is the API version of that idea: the provider keeps a processed prefix between requests and bills it at a discount.
KV cache vs context window
The context window is the most tokens a model can handle. The KV cache is the memory the tokens you actually use take up, which is why local runtimes let you set a shorter context to save memory.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary