AI glossary · Local AI
What is LLM quantisation?
Also called: quantization, LLM quantization, model quantisation
Definition
Quantisation is storing a model’s weights in fewer bits than the 16 per weight most models are released with, so the model needs less memory and runs on smaller hardware, at a small cost in accuracy.
Explained
How it works
Most open models are released with each weight as a 16-bit number (FP16 or BF16). Quantisation maps the weights onto far fewer levels. llama.cpp’s formats split each tensor into small blocks that share a scale factor, so a weight needs only a few bits plus its share of the scales. The Q4_K type itself uses 4.5 bits per weight, and Q4_K_M, which keeps some tensors at 6 bits, averages 4.89 on Llama 3.1 8B.
The names describe the recipe. Q8_0 is 8-bit. The _K types (k-quants) group blocks into super-blocks, and the S, M and L variants keep progressively more of the sensitive tensors at higher precision: Q4_K_M stores half of the attention-value and feed-forward output tensors at 6 bits and the rest at 4. These files use the GGUF format.
Quantisation shrinks the weights only. The KV cache that holds your context is separate, with its own optional quantisation setting.
Example
Llama 3.1 8B at five precisions
With an 8K context, Llama 3.1 8B needs 18 GB at FP16 but 6.6 GB at Q4_K_M, small enough for an 8 GB card such as the GeForce RTX 4060. The quality cost, measured by llama.cpp’s maintainers on the closely related Llama 3 8B, is tiny at Q8_0, small at Q4_K_M and large at Q2_K.
| Format | Bits per weight | Weights | Total at 8K | Perplexity vs FP16 |
|---|---|---|---|---|
| FP16 / BF16 | 16 | 15.0 GB | 17.6 GB | Baseline |
| Q8_0 (8-bit) | 8.5 | 7.9 GB | 9.9 GB | +0.04% |
| Q6_K | 6.56 | 6.1 GB | 8.1 GB | +0.35% |
| Q4_K_M | 4.89 | 4.6 GB | 6.6 GB | +2.8% |
| Q2_K | 3.16 | 3.0 GB | 5.0 GB | +56.5% |
Memory from our VRAM calculator: parameters × llama.cpp’s bits per weight, plus an FP16 KV cache and overhead, batch 1 (GB = 1,024³ bytes). Perplexity from the llama.cpp perplexity README (Llama 3 8B, Wikitext-2, no importance matrix): how well the model predicts text, lower is better.
Cost and quality
Why it matters
Quantisation decides which models you can run at all. Q4_K_M weights are about 31% of the size of the 16-bit original, which moves a model down whole tiers of hardware. Use our VRAM calculator to see what each level needs with your context length.
Q4_K_M is a common starting point; step up to Q5_K_M, Q6_K or Q8_0 if they fit. The loss differs between models, and perplexity is a proxy rather than a task score, so check quality on your own prompts.
Don’t mix up
Common confusions
- A quantised model vs a smaller model
- A quantised 14B model still has 14 billion parameters and the same layers, stored more compactly. A smaller model has fewer parameters. The two save memory in different ways and lose quality in different ways.
- Weight quantisation vs KV cache quantisation
- Q4_K_M in a file name describes the weights. KV cache quantisation (
q8_0in llama.cpp and Ollama) is a separate runtime setting that roughly halves the memory your context uses.
Go deeper
Try it and read more
- Free toolGPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.
- Guide · 12 min readHow much VRAM do you need to run an LLM locally?VRAM needed for 8B to 70B models at Q4, Q8 and FP16, what fits on 8 to 80 GB GPUs, and how context length, KV cache and CPU offload change it.
- Guide · 10 min readOllama vs LM Studio vs llama.cpp: which local LLM tool?Ollama vs LM Studio vs llama.cpp from their own docs: install, GGUF and MLX, APIs and default ports, GPU support, context defaults, licences and which to pick.
Related
Related terms
- GGUFGGUF is a single-file binary format from the ggml and llama.cpp project that stores a model’s weights together with everything needed to run them, such as its architecture, tokenizer and quantisation type.
- VRAMVRAM is the memory on a graphics card, and for running AI models locally it is the main limit: a model runs at full GPU speed only when its weights, KV cache and working buffers all fit in it.
- KV cacheThe KV cache is the memory where a language model keeps the attention keys and values of every token it has already processed, so each new token is computed without reprocessing the whole sequence.
- Open weightsOpen weights means a model’s trained parameters are published for anyone to download and run on their own hardware, under a licence that sets what they may do with them.
- LoRALoRA (low-rank adaptation) is a fine-tuning method that freezes a model’s original weights and trains two small matrices per adapted layer, whose product is added to the frozen weights, so only a tiny fraction of parameters is trained.