Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Local AI

What is LLM quantisation?

Also called: quantization, LLM quantization, model quantisation

Definition

Quantisation is storing a model’s weights in fewer bits than the 16 per weight most models are released with, so the model needs less memory and runs on smaller hardware, at a small cost in accuracy.

Explained

How it works

Most open models are released with each weight as a 16-bit number (FP16 or BF16). Quantisation maps the weights onto far fewer levels. llama.cpp’s formats split each tensor into small blocks that share a scale factor, so a weight needs only a few bits plus its share of the scales. The Q4_K type itself uses 4.5 bits per weight, and Q4_K_M, which keeps some tensors at 6 bits, averages 4.89 on Llama 3.1 8B.

The names describe the recipe. Q8_0 is 8-bit. The _K types (k-quants) group blocks into super-blocks, and the S, M and L variants keep progressively more of the sensitive tensors at higher precision: Q4_K_M stores half of the attention-value and feed-forward output tensors at 6 bits and the rest at 4. These files use the GGUF format.

Quantisation shrinks the weights only. The KV cache that holds your context is separate, with its own optional quantisation setting.

Example

Llama 3.1 8B at five precisions

With an 8K context, Llama 3.1 8B needs 18 GB at FP16 but 6.6 GB at Q4_K_M, small enough for an 8 GB card such as the GeForce RTX 4060. The quality cost, measured by llama.cpp’s maintainers on the closely related Llama 3 8B, is tiny at Q8_0, small at Q4_K_M and large at Q2_K.

Llama 3.1 8B memory, and quality loss measured on Llama 3 8B
FormatBits per weightWeightsTotal at 8KPerplexity vs FP16
FP16 / BF161615.0 GB17.6 GBBaseline
Q8_0 (8-bit)8.57.9 GB9.9 GB+0.04%
Q6_K6.566.1 GB8.1 GB+0.35%
Q4_K_M4.894.6 GB6.6 GB+2.8%
Q2_K3.163.0 GB5.0 GB+56.5%

Memory from our VRAM calculator: parameters × llama.cpp’s bits per weight, plus an FP16 KV cache and overhead, batch 1 (GB = 1,024³ bytes). Perplexity from the llama.cpp perplexity README (Llama 3 8B, Wikitext-2, no importance matrix): how well the model predicts text, lower is better.

Cost and quality

Why it matters

Quantisation decides which models you can run at all. Q4_K_M weights are about 31% of the size of the 16-bit original, which moves a model down whole tiers of hardware. Use our VRAM calculator to see what each level needs with your context length.

Q4_K_M is a common starting point; step up to Q5_K_M, Q6_K or Q8_0 if they fit. The loss differs between models, and perplexity is a proxy rather than a task score, so check quality on your own prompts.

Don’t mix up

Common confusions

A quantised model vs a smaller model
A quantised 14B model still has 14 billion parameters and the same layers, stored more compactly. A smaller model has fewer parameters. The two save memory in different ways and lose quality in different ways.
Weight quantisation vs KV cache quantisation
Q4_K_M in a file name describes the weights. KV cache quantisation (q8_0 in llama.cpp and Ollama) is a separate runtime setting that roughly halves the memory your context uses.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary