AI glossary · Local AI
What are open weights?
Also called: open-weight model, downloadable model weights
Definition
Open weights means a model’s trained parameters are published for anyone to download and run on their own hardware, under a licence that sets what they may do with them.
Explained
How it works
A trained model is an architecture plus billions of learned numbers, its weights. An open-weight release publishes those numbers, usually on Hugging Face with a config file and tokenizer, so you can run the model yourself, quantise it, fine-tune it or host it for others. A closed model is only available as a hosted service, through its maker’s API or the cloud platforms it licenses.
Open weights is not the same as open source. The Open Source Initiative’s Open Source AI Definition 1.0 asks for three things under OSI-approved terms: the parameters, the complete code used to train and run the system, and sufficiently detailed information about the training data. Many open-weight releases publish the weights and the code to run them, but not the full training code or data information.
The licence decides what you may do. gpt-oss and Qwen3 30B-A3B use Apache 2.0 and DeepSeek R1 uses MIT. Llama 3.3 uses Meta’s own Llama 3.3 Community License Agreement, which asks redistributors to display “Built with Llama” and requires any company whose products had more than 700 million monthly active users on the Llama 3.3 release date to request a separate licence from Meta.
Example
Four open-weight models, four sets of terms
In our daily price data, 146 of 335 models link to downloadable weights on Hugging Face, so the same model is often sold through several APIs and also free to run yourself. Here are four, with the licence from each model card and the memory each needs on your own machine.
| Model | Licence | Download | Memory at 8K context |
|---|---|---|---|
| gpt-oss 20B | Apache 2.0 | Open | 13 GB |
| Qwen3 30B-A3B | Apache 2.0 | Open | 20 GB |
| DeepSeek R1 (671B) | MIT | Open | 421 GB |
| Llama 3.3 70B | Llama 3.3 Community License | Gated: share your contact details first | 47 GB |
Licences from the Hugging Face model cards, checked 2026-10-11. Price data 2026-10-11. Memory from our VRAM calculator: Q4_K_M weights (gpt-oss 20B uses its official MXFP4 file), FP16 KV cache, batch 1, plus overhead; model configs fetched 2026-10-08. GB = 1,024³ bytes.
Cost and quality
Why it matters
Open weights give you options an API can’t: run offline, keep data on your own machines, pin one exact version for as long as you like, fine-tune, or compare hosting providers for the same model. The trade-off is that you supply the hardware, and the VRAM you need grows with the model.
Read the licence before you ship. “Open” on a model page can mean anything from Apache 2.0 to custom terms with attribution rules and user thresholds.
Don’t mix up
Common confusions
- Open weights vs open source
- Open source AI, in the OSI’s definition, also needs the training code and detailed information about the training data, so others can study and rebuild the model. A release with weights alone is open weights, whatever its announcement calls it.
- Free to download vs free of conditions
- Custom licences such as Llama’s add attribution and user-count terms, and some downloads are gated until you accept them. Apache 2.0 and MIT are standard licences with no user-count or branding rules; their main conditions are to include the licence and keep the copyright notices, and Apache 2.0 also asks you to mark files you change.
Go deeper
Try it and read more
- Free toolGPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.
- Free toolAI Model ComparisonCompare prices, context windows and features across models.
- Guide · 12 min readHow much VRAM do you need to run an LLM locally?VRAM needed for 8B to 70B models at Q4, Q8 and FP16, what fits on 8 to 80 GB GPUs, and how context length, KV cache and CPU offload change it.
- Guide · 10 min readOllama vs LM Studio vs llama.cpp: which local LLM tool?Ollama vs LM Studio vs llama.cpp from their own docs: install, GGUF and MLX, APIs and default ports, GPU support, context defaults, licences and which to pick.
Related
Related terms
- QuantisationQuantisation is storing a model’s weights in fewer bits than the 16 per weight most models are released with, so the model needs less memory and runs on smaller hardware, at a small cost in accuracy.
- GGUFGGUF is a single-file binary format from the ggml and llama.cpp project that stores a model’s weights together with everything needed to run them, such as its architecture, tokenizer and quantisation type.
- VRAMVRAM is the memory on a graphics card, and for running AI models locally it is the main limit: a model runs at full GPU speed only when its weights, KV cache and working buffers all fit in it.
- Fine-tuningFine-tuning is training an existing model further on a set of your own example inputs and ideal outputs, so its weights change and it follows a task, format or style without long instructions in every prompt.
- LoRALoRA (low-rank adaptation) is a fine-tuning method that freezes a model’s original weights and trains two small matrices per adapted layer, whose product is added to the frozen weights, so only a tiny fraction of parameters is trained.