Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Data and training

What is LoRA (low-rank adaptation)?

Also called: low-rank adaptation, LoRA adapter, LoRA fine-tuning

Definition

LoRA (low-rank adaptation) is a fine-tuning method that freezes a model’s original weights and trains two small matrices per adapted layer, whose product is added to the frozen weights, so only a tiny fraction of parameters is trained.

Explained

How it works

A layer’s weight matrix W₀ has d × k entries. LoRA, introduced in a 2021 paper (arXiv 2106.09685), leaves W₀ frozen and learns the change as the product of two thin matrices: B, with d × r entries, and A, with r × k, where the rank r is much smaller than d or k. The layer computes W₀x + BAx. B starts at zero, so training begins from the unchanged model.

Only A and B receive gradients, so the memory-hungry optimiser state is kept for a tiny fraction of the parameters. After training you can merge BA into W₀, which the paper says adds no inference latency, or keep the adapter separate and swap adapters on one base model.

The paper’s abstract reports that, compared with fine-tuning GPT-3 175B with Adam, LoRA can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times, with quality on par with or better than full fine-tuning on the models it tested.

Example

Rank-8 LoRA on Llama 3.1 8B

Take one 4,096 × 4,096 attention projection. Full fine-tuning updates 16,777,216 weights; LoRA at r = 8 trains 8 × (4,096 + 4,096) = 65,536, or 0.39%.

Apply it as the paper’s experiments did, to the query and value projections, across all 32 layers. The value projection is 4,096 × 1,024, because the model uses 8 key-value heads for its 32 query heads (grouped-query attention). That trains 3,407,872 parameters, 0.042% of the model’s 8,030,261,248, and the adapter is about 6.8 MB in 16-bit precision.

Trainable parameters per layer at r = 8
MatrixShapeFull fine-tuningLoRA
Query projection4,096 × 4,09616,777,21665,536
Value projection4,096 × 1,0244,194,30440,960
All 32 layers671,088,6403,407,872

Hidden size from the model’s config.json; layer count, value width and the 8,030,261,248-parameter total from our VRAM model data. LoRA parameters = r × (d + k).

Cost and quality

Why it matters

LoRA is why tuning an open-weights model can fit on one GPU: the base weights load frozen, even quantised, and only the small adapter needs gradients and optimiser state. Adapters are megabytes, so you can keep one per task or customer on a single base model.

Rank is the main setting: each step up in r adds d + k parameters per adapted matrix. To run a tuned model you still need memory for the full base model, which the VRAM calculator estimates.

Don’t mix up

Common confusions

LoRA vs QLoRA
QLoRA (arXiv 2305.14314) trains LoRA adapters on top of a base model quantised to 4 bits, using a 4-bit NormalFloat data type. Its abstract reports fine-tuning a 65B-parameter model on a single 48 GB GPU.
LoRA vs full fine-tuning
Full fine-tuning updates every weight and produces a full-size copy of the model. LoRA produces a small adapter that only works with the exact base model it was trained on.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary