Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Data and training

What is fine-tuning an LLM?

Also called: finetuning, supervised fine-tuning, SFT, model tuning

Definition

Fine-tuning is training an existing model further on a set of your own example inputs and ideal outputs, so its weights change and it follows a task, format or style without long instructions in every prompt.

Explained

How it works

You prepare examples, usually as JSONL (one JSON conversation per line) ending with the answer you want, and run a training job. Supervised fine-tuning teaches the model to produce those answers; other variants train on a preferred and a rejected answer, or on scores from a grader. Full fine-tuning updates every weight, while parameter-efficient methods such as LoRA freeze the original weights and train a small add-on.

Hosted tuning has narrowed. OpenAI is winding down self-serve fine-tuning: organisations that had never fine-tuned lost access on 7 May 2026, from 2 July 2026 only organisations that had used a fine-tuned model in the previous 60 days could start new jobs, and those remaining customers can create jobs until 6 January 2027. The Gemini API says no model there supports fine-tuning and points to Google Cloud’s Gemini Enterprise Agent Platform instead. Mistral marks its fine-tuning feature deprecated. With open-weights models you can still tune on your own hardware.

Example

A ticket classifier: examples in the prompt or in the weights

You want support tickets sorted into four labels. One labelled example, “Ticket: I was charged twice for my March invoice. Label: billing”, is 14 tokens in o200k_base. Few-shot prompting with 20 such examples adds about 280 tokens to every request: on GPT-4.1 Mini at $0.40 per million input tokens, that is $11.20 per 100,000 requests.

Fine-tuning moves those examples into the weights, but you pay for training (billed on the tokens in your file times the number of passes over it), for building and checking a dataset, and, at OpenAI, a higher per-token price to use the tuned model. When the prompt version is this cheap and works, tuning rarely pays. It earns its keep when prompts grow long, results stay inconsistent, or a small tuned model can replace a larger one.

Cost and quality

Why it matters

Fine-tuning can make a small, cheap model reliable at one narrow job and keep output formats consistent with short prompts. It is a poor way to keep facts current, because every change means another training run; RAG handles changing knowledge better.

It also creates lock-in. A tuned model exists only on the platform and base model you trained it on, and OpenAI says inference on fine-tuned models stops when the underlying base model is deprecated. Keep your dataset: it doubles as an evaluation set for whichever model you move to.

Don’t mix up

Common confusions

Fine-tuning vs RAG
Fine-tuning changes how a model behaves; RAG changes what it reads at question time. Use RAG for knowledge, fine-tuning for task, tone and format. Many systems use both.
Fine-tuning vs prompting
A system prompt with examples changes behaviour per request and can be edited in seconds. Fine-tuning bakes behaviour in and needs a new training run to change. Try prompting first.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary

Written by Tahir Nazir. Checked .

How this was checked: Availability checked on 2026-10-11 against OpenAI’s deprecations page (7 May 2026, 2 July 2026 and 6 January 2027 dates; base-model deprecation rule), the Gemini API model-tuning page and Mistral’s fine-tuning docs; the higher price for using a tuned model checked on OpenAI’s pricing page. The example’s token count was measured with gpt-tokenizer 4.0.0 and its price comes from our daily data (2026-10-11). No training jobs were run.