AI glossary · Data and training
What is LoRA (low-rank adaptation)?
Also called: low-rank adaptation, LoRA adapter, LoRA fine-tuning
Definition
LoRA (low-rank adaptation) is a fine-tuning method that freezes a model’s original weights and trains two small matrices per adapted layer, whose product is added to the frozen weights, so only a tiny fraction of parameters is trained.
Explained
How it works
A layer’s weight matrix W₀ has d × k entries. LoRA, introduced in a 2021 paper (arXiv 2106.09685), leaves W₀ frozen and learns the change as the product of two thin matrices: B, with d × r entries, and A, with r × k, where the rank r is much smaller than d or k. The layer computes W₀x + BAx. B starts at zero, so training begins from the unchanged model.
Only A and B receive gradients, so the memory-hungry optimiser state is kept for a tiny fraction of the parameters. After training you can merge BA into W₀, which the paper says adds no inference latency, or keep the adapter separate and swap adapters on one base model.
The paper’s abstract reports that, compared with fine-tuning GPT-3 175B with Adam, LoRA can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times, with quality on par with or better than full fine-tuning on the models it tested.
Example
Rank-8 LoRA on Llama 3.1 8B
Take one 4,096 × 4,096 attention projection. Full fine-tuning updates 16,777,216 weights; LoRA at r = 8 trains 8 × (4,096 + 4,096) = 65,536, or 0.39%.
Apply it as the paper’s experiments did, to the query and value projections, across all 32 layers. The value projection is 4,096 × 1,024, because the model uses 8 key-value heads for its 32 query heads (grouped-query attention). That trains 3,407,872 parameters, 0.042% of the model’s 8,030,261,248, and the adapter is about 6.8 MB in 16-bit precision.
| Matrix | Shape | Full fine-tuning | LoRA |
|---|---|---|---|
| Query projection | 4,096 × 4,096 | 16,777,216 | 65,536 |
| Value projection | 4,096 × 1,024 | 4,194,304 | 40,960 |
| All 32 layers | 671,088,640 | 3,407,872 |
Hidden size from the model’s config.json; layer count, value width and the 8,030,261,248-parameter total from our VRAM model data. LoRA parameters = r × (d + k).
Cost and quality
Why it matters
LoRA is why tuning an open-weights model can fit on one GPU: the base weights load frozen, even quantised, and only the small adapter needs gradients and optimiser state. Adapters are megabytes, so you can keep one per task or customer on a single base model.
Rank is the main setting: each step up in r adds d + k parameters per adapted matrix. To run a tuned model you still need memory for the full base model, which the VRAM calculator estimates.
Don’t mix up
Common confusions
- LoRA vs QLoRA
- QLoRA (arXiv 2305.14314) trains LoRA adapters on top of a base model quantised to 4 bits, using a 4-bit NormalFloat data type. Its abstract reports fine-tuning a 65B-parameter model on a single 48 GB GPU.
- LoRA vs full fine-tuning
- Full fine-tuning updates every weight and produces a full-size copy of the model. LoRA produces a small adapter that only works with the exact base model it was trained on.
Go deeper
Try it and read more
- Free toolGPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.
- Free toolLlama 3.1 8B VRAM requirementsLlama 3.1 8B needs about 7 GB of VRAM at Q4_K_M and 10 GB at 8-bit with an 8K context.
- Guide · 13 min readFine-tuning JSONL format for OpenAI, Gemini and MistralThe exact fine-tuning JSONL format for OpenAI, Gemini and Mistral, with a tested validator, CSV to JSONL scripts, common errors and training cost maths.
- Guide · 12 min readHow much VRAM do you need to run an LLM locally?VRAM needed for 8B to 70B models at Q4, Q8 and FP16, what fits on 8 to 80 GB GPUs, and how context length, KV cache and CPU offload change it.
Related
Related terms
- Fine-tuningFine-tuning is training an existing model further on a set of your own example inputs and ideal outputs, so its weights change and it follows a task, format or style without long instructions in every prompt.
- QuantisationQuantisation is storing a model’s weights in fewer bits than the 16 per weight most models are released with, so the model needs less memory and runs on smaller hardware, at a small cost in accuracy.
- VRAMVRAM is the memory on a graphics card, and for running AI models locally it is the main limit: a model runs at full GPU speed only when its weights, KV cache and working buffers all fit in it.
- Open weightsOpen weights means a model’s trained parameters are published for anyone to download and run on their own hardware, under a licence that sets what they may do with them.