Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Models and context

What is few-shot prompting?

Also called: multishot prompting, in-context examples, k-shot prompting

Definition

Few-shot prompting is giving a model a handful of worked input-and-output examples inside the prompt, so it copies the pattern for a new input without any retraining.

Explained

How it works

A “shot” is an example. Zero-shot gives only the instruction, one-shot adds one example, and few-shot adds several. The model reads the examples as part of its context window and infers the format, labels and tone from them. Nothing about the model changes: remove the examples and the behaviour goes with them.

OpenAI’s 2020 GPT-3 paper, “Language Models are Few-Shot Learners”, showed a 175-billion-parameter model handling translation, question answering and 3-digit arithmetic from demonstrations given purely as text, with no gradient updates or fine-tuning.

Current advice: Anthropic suggests 3 to 5 relevant, varied examples wrapped in <example> tags. For reasoning models, OpenAI says to try zero-shot first and add examples only if needed, keeping them closely aligned with your instructions.

Example

Three examples for a ticket classifier

The zero-shot version of this prompt is 44 tokens with OpenAI’s o200k_base tokenizer. Adding the three labelled examples below makes it 90, 46 more per request. On GPT-6 Luna at $0.10 per million input tokens, that is $4.60 for every 1,000,000 requests.

That is cheap insurance for a classifier that might otherwise invent labels or explain itself. Examples get expensive when they are long, such as whole documents or code files sent on every call. Keep them unchanged at the start of the prompt so prompt caching can reuse them once the prefix is long enough.

Few-shot prompt (90 tokens, o200k_base)
Classify the customer message as one of: billing, delivery, product, other. Reply with the label only.

Message: I was charged twice for the same order.
Label: billing

Message: The left brake lever snapped after two rides.
Label: product

Message: Can I change the delivery address on my order?
Label: delivery

Message: My parcel was meant to arrive on Tuesday and there is still no sign of it.
Label:

Cost and quality

Why it matters

Examples are often the quickest fix for format and label drift, and they show edge cases better than a paragraph of rules. Anthropic calls them one of the most reliable ways to steer output format, tone and structure.

But the model copies everything, including accidents: three examples with short answers teach short answers. Vary them deliberately, and compare results with and without them on your own test set rather than assuming they help.

Don’t mix up

Common confusions

Few-shot prompting vs fine-tuning
Few-shot examples live in the prompt and are paid for on every request. Fine-tuning trains them into the model’s weights once, so they no longer need sending, at the cost of a training job and a custom model.
Few-shot vs chain-of-thought
Plain few-shot examples show the answer. Chain-of-thought examples also show the working; the original chain-of-thought paper used few-shot examples that included reasoning steps.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary