AI glossary · Data and training
What is fine-tuning an LLM?
Also called: finetuning, supervised fine-tuning, SFT, model tuning
Definition
Fine-tuning is training an existing model further on a set of your own example inputs and ideal outputs, so its weights change and it follows a task, format or style without long instructions in every prompt.
Explained
How it works
You prepare examples, usually as JSONL (one JSON conversation per line) ending with the answer you want, and run a training job. Supervised fine-tuning teaches the model to produce those answers; other variants train on a preferred and a rejected answer, or on scores from a grader. Full fine-tuning updates every weight, while parameter-efficient methods such as LoRA freeze the original weights and train a small add-on.
Hosted tuning has narrowed. OpenAI is winding down self-serve fine-tuning: organisations that had never fine-tuned lost access on 7 May 2026, from 2 July 2026 only organisations that had used a fine-tuned model in the previous 60 days could start new jobs, and those remaining customers can create jobs until 6 January 2027. The Gemini API says no model there supports fine-tuning and points to Google Cloud’s Gemini Enterprise Agent Platform instead. Mistral marks its fine-tuning feature deprecated. With open-weights models you can still tune on your own hardware.
Example
A ticket classifier: examples in the prompt or in the weights
You want support tickets sorted into four labels. One labelled example, “Ticket: I was charged twice for my March invoice. Label: billing”, is 14 tokens in o200k_base. Few-shot prompting with 20 such examples adds about 280 tokens to every request: on GPT-4.1 Mini at $0.40 per million input tokens, that is $11.20 per 100,000 requests.
Fine-tuning moves those examples into the weights, but you pay for training (billed on the tokens in your file times the number of passes over it), for building and checking a dataset, and, at OpenAI, a higher per-token price to use the tuned model. When the prompt version is this cheap and works, tuning rarely pays. It earns its keep when prompts grow long, results stay inconsistent, or a small tuned model can replace a larger one.
Cost and quality
Why it matters
Fine-tuning can make a small, cheap model reliable at one narrow job and keep output formats consistent with short prompts. It is a poor way to keep facts current, because every change means another training run; RAG handles changing knowledge better.
It also creates lock-in. A tuned model exists only on the platform and base model you trained it on, and OpenAI says inference on fine-tuned models stops when the underlying base model is deprecated. Keep your dataset: it doubles as an evaluation set for whichever model you move to.
Don’t mix up
Common confusions
- Fine-tuning vs RAG
- Fine-tuning changes how a model behaves; RAG changes what it reads at question time. Use RAG for knowledge, fine-tuning for task, tone and format. Many systems use both.
- Fine-tuning vs prompting
- A system prompt with examples changes behaviour per request and can be edited in seconds. Fine-tuning bakes behaviour in and needs a new training run to change. Try prompting first.
Go deeper
Try it and read more
- Free toolAI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Guide · 13 min readFine-tuning JSONL format for OpenAI, Gemini and MistralThe exact fine-tuning JSONL format for OpenAI, Gemini and Mistral, with a tested validator, CSV to JSONL scripts, common errors and training cost maths.
Related
Related terms
- LoRALoRA (low-rank adaptation) is a fine-tuning method that freezes a model’s original weights and trains two small matrices per adapted layer, whose product is added to the frozen weights, so only a tiny fraction of parameters is trained.
- RAGRAG (retrieval-augmented generation) is a technique where an application searches your own documents for passages relevant to a question and adds them to the prompt, so the model answers from that text rather than from memory alone.
- Few-shot promptingFew-shot prompting is giving a model a handful of worked input-and-output examples inside the prompt, so it copies the pattern for a new input without any retraining.
- Open weightsOpen weights means a model’s trained parameters are published for anyone to download and run on their own hardware, under a licence that sets what they may do with them.
- System promptA system prompt is the set of standing instructions a developer sends with every request to set a model’s role, rules and output format, kept separate from what the user types.