AI glossary · Local AI
What is a GGUF file?
Also called: GGUF file, .gguf, GGUF format
Definition
GGUF is a single-file binary format from the ggml and llama.cpp project that stores a model’s weights together with everything needed to run them, such as its architecture, tokenizer and quantisation type.
Explained
How it works
GGUF replaced the project’s earlier GGML, GGMF and GGJT formats. A file starts with the magic bytes GGUF and a version number (currently 3), then metadata as key-value pairs (architecture, context length, tokenizer and so on), then a description of every tensor, then the tensor data itself. The layout can be memory-mapped, so models load quickly.
Models are usually trained in PyTorch and published as safetensors files. llama.cpp’s convert_hf_to_gguf.py turns them into a 16-bit GGUF, and llama-quantize makes the smaller quantised versions from it.
The file name ends with the encoding, so Qwen3-14B-Q4_K_M.gguf is Qwen3 14B at Q4_K_M. Very large models are split into shards named like -00001-of-00003.gguf. llama.cpp reads GGUF directly, LM Studio runs GGUF models with llama.cpp, and Ollama imports a GGUF file through a Modelfile with a FROM line.
Example
Qwen’s official GGUF files for Qwen3 14B
Qwen publishes Qwen3 14B as GGUF files at several quantisations, one file per level. We read their sizes from the Hugging Face API and compared them with our VRAM calculator’s weight estimate: every one is within 0.3%. The file is only the weights. Running it with an 8K context adds the KV cache and working space for the runtime.
| File | File size | Our weight estimate | Total at 8K context |
|---|---|---|---|
Qwen3-14B-Q4_K_M.gguf | 8.4 GB | 8.4 GB | 10.7 GB |
Qwen3-14B-Q5_K_M.gguf | 9.8 GB | 9.8 GB | 12.2 GB |
Qwen3-14B-Q6_K.gguf | 11.3 GB | 11.3 GB | 13.8 GB |
Qwen3-14B-Q8_0.gguf | 14.6 GB | 14.6 GB | 17.5 GB |
File sizes from the Hugging Face API (Qwen/Qwen3-14B-GGUF), 2026-10-11. Estimates: parameters × llama.cpp’s bits per weight; totals add an FP16 KV cache and overhead, batch 1. Model config fetched 2026-10-08. GB = 1,024³ bytes.
Cost and quality
Why it matters
GGUF is what most local tools download, and its name tells you the memory you need before you download it: the quantisation sets the file size, and the context you want adds the rest. Check the total against your VRAM.
One self-contained file also moves easily between tools: the same GGUF runs in llama.cpp and LM Studio, and Ollama can import it.
Don’t mix up
Common confusions
- GGUF vs safetensors
- Safetensors is the Hugging Face weights format that training frameworks save to, with the config and tokenizer in separate files. GGUF bundles all of it in one file, usually quantised, for runtimes built on ggml.
- GGUF vs GGML
- ggml is the tensor library that llama.cpp is built on. GGML was also the name of its older file format, which GGUF replaced, so a file described as a “GGML model” is the old format, not the library.
Go deeper
Try it and read more
- Free toolGPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.
- Guide · 10 min readOllama vs LM Studio vs llama.cpp: which local LLM tool?Ollama vs LM Studio vs llama.cpp from their own docs: install, GGUF and MLX, APIs and default ports, GPU support, context defaults, licences and which to pick.
- Guide · 12 min readHow much VRAM do you need to run an LLM locally?VRAM needed for 8B to 70B models at Q4, Q8 and FP16, what fits on 8 to 80 GB GPUs, and how context length, KV cache and CPU offload change it.
Related
Related terms
- QuantisationQuantisation is storing a model’s weights in fewer bits than the 16 per weight most models are released with, so the model needs less memory and runs on smaller hardware, at a small cost in accuracy.
- VRAMVRAM is the memory on a graphics card, and for running AI models locally it is the main limit: a model runs at full GPU speed only when its weights, KV cache and working buffers all fit in it.
- Open weightsOpen weights means a model’s trained parameters are published for anyone to download and run on their own hardware, under a licence that sets what they may do with them.
- KV cacheThe KV cache is the memory where a language model keeps the attention keys and values of every token it has already processed, so each new token is computed without reprocessing the whole sequence.