AI glossary · Models and context
What is few-shot prompting?
Also called: multishot prompting, in-context examples, k-shot prompting
Definition
Few-shot prompting is giving a model a handful of worked input-and-output examples inside the prompt, so it copies the pattern for a new input without any retraining.
Explained
How it works
A “shot” is an example. Zero-shot gives only the instruction, one-shot adds one example, and few-shot adds several. The model reads the examples as part of its context window and infers the format, labels and tone from them. Nothing about the model changes: remove the examples and the behaviour goes with them.
OpenAI’s 2020 GPT-3 paper, “Language Models are Few-Shot Learners”, showed a 175-billion-parameter model handling translation, question answering and 3-digit arithmetic from demonstrations given purely as text, with no gradient updates or fine-tuning.
Current advice: Anthropic suggests 3 to 5 relevant, varied examples wrapped in <example> tags. For reasoning models, OpenAI says to try zero-shot first and add examples only if needed, keeping them closely aligned with your instructions.
Example
Three examples for a ticket classifier
The zero-shot version of this prompt is 44 tokens with OpenAI’s o200k_base tokenizer. Adding the three labelled examples below makes it 90, 46 more per request. On GPT-6 Luna at $0.10 per million input tokens, that is $4.60 for every 1,000,000 requests.
That is cheap insurance for a classifier that might otherwise invent labels or explain itself. Examples get expensive when they are long, such as whole documents or code files sent on every call. Keep them unchanged at the start of the prompt so prompt caching can reuse them once the prefix is long enough.
Classify the customer message as one of: billing, delivery, product, other. Reply with the label only.
Message: I was charged twice for the same order.
Label: billing
Message: The left brake lever snapped after two rides.
Label: product
Message: Can I change the delivery address on my order?
Label: delivery
Message: My parcel was meant to arrive on Tuesday and there is still no sign of it.
Label:Cost and quality
Why it matters
Examples are often the quickest fix for format and label drift, and they show edge cases better than a paragraph of rules. Anthropic calls them one of the most reliable ways to steer output format, tone and structure.
But the model copies everything, including accidents: three examples with short answers teach short answers. Vary them deliberately, and compare results with and without them on your own test set rather than assuming they help.
Don’t mix up
Common confusions
- Few-shot prompting vs fine-tuning
- Few-shot examples live in the prompt and are paid for on every request. Fine-tuning trains them into the model’s weights once, so they no longer need sending, at the cost of a training job and a custom model.
- Few-shot vs chain-of-thought
- Plain few-shot examples show the answer. Chain-of-thought examples also show the working; the original chain-of-thought paper used few-shot examples that included reasoning steps.
Go deeper
Try it and read more
- Free toolAI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Guide · 11 min readHow to estimate LLM API costs: the formula, worked examples and the trapsEstimate LLM API costs per request and per month: the token formula, worked examples for chatbots, RAG and agents, plus caching, batch and reasoning tokens.
Related
Related terms
- Chain of thoughtChain-of-thought prompting is asking a language model to write out intermediate reasoning steps before its final answer, usually by showing worked examples or telling it to think step by step.
- Fine-tuningFine-tuning is training an existing model further on a set of your own example inputs and ideal outputs, so its weights change and it follows a task, format or style without long instructions in every prompt.
- System promptA system prompt is the set of standing instructions a developer sends with every request to set a model’s role, rules and output format, kept separate from what the user types.
- Prompt cachingPrompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.