AI glossary · Tokens and cost
What is prompt caching?
Also called: context caching
Definition
Prompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.
Explained
How it works
Most requests to a model repeat a long opening: the system prompt, tool definitions, a document, or the conversation so far. With caching on, the provider keeps the model’s work on that opening (the prefix) for a few minutes. If the next request starts with the identical tokens, it reads the prefix from the cache instead of processing it again.
Only an exact prefix counts. One changed character early in the prompt, such as a timestamp in the system prompt, means everything after it is new. OpenAI and Gemini cache automatically above a minimum length; Anthropic caches what you mark with cache_control, or the last cacheable block when you turn on automatic caching. Caches expire after a short idle time (Anthropic’s default is 5 minutes, refreshed on every hit), and some providers charge extra for the first write.
Example
A 20,000-token prefix reused 100 times
A support bot sends a 20,000-token system prompt and help-centre extract with every question, then a 300-token question, and gets a 400-token answer. On Claude Sonnet 5.5, input costs $2 per million tokens, a cache write $2.50 and a cache read $0.10.
Without caching, 100 questions cost $4.46. With caching, the first request writes the prefix and the other 99 read it, for $0.71 in total: 84% less, from one setting. The answers cost the same either way, because caching only discounts input.
| Without caching | With caching | |
|---|---|---|
| Prefix tokens billed at full price | 2,000,000 | 0 |
| Prefix tokens written to cache | 0 | 20,000 |
| Prefix tokens read from cache | 0 | 1,980,000 |
| Total cost | $4.46 | $0.71 |
Prices from our daily data, 2026-10-11. Assumes the requests arrive often enough to keep the cache warm.
Cost and quality
Why it matters
Input is usually most of an agent’s or chatbot’s bill, because the same context is sent on every turn. Caching is often the single largest saving available without changing the model, and it also cuts the time to the first token on long prompts.
It only pays when the prefix is reused before it expires. Slow traffic, or a prompt that changes near the top, can mean paying the write premium with no reads to recover it.
Don’t mix up
Common confusions
- Prompt caching vs response caching
- A response cache in your own code returns a stored answer to an identical question without calling the model. Prompt caching still calls the model and generates a fresh answer; only the processing of the repeated input is reused.
- Prompt caching vs the KV cache
- The KV cache is the memory a model keeps during a single request. Prompt caching keeps that work between requests, on the provider’s side, and bills it at a discount.
Go deeper
Try it and read more
- Free toolPrompt Caching CalculatorEstimate savings from prompt caching.
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Guide · 11 min readPrompt caching explained: how it works on OpenAI, Anthropic and Gemini, and when it paysHow prompt caching works on OpenAI, Anthropic and Gemini: what gets cached, minimum lengths, how long caches last, write and read prices, and when it pays off.
- Guide · 11 min readHow to reduce Claude Code token usage, and what each fix costs youCut Claude Code token use with /clear, /compact, model and effort choice, a lean CLAUDE.md, fewer MCP servers and subagents, and the trade-off of each.
Related
Related terms
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- KV cacheThe KV cache is the memory where a language model keeps the attention keys and values of every token it has already processed, so each new token is computed without reprocessing the whole sequence.
- Batch APIA batch API is an asynchronous way to send many model requests as one job, which the provider works through when it has capacity, typically within 24 hours, at half the normal per-token price on OpenAI, Anthropic and Google.
- Context windowA context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.
- Time to first tokenTime to first token (TTFT) is how long a model takes from receiving a request to returning the first token of its reply, the wait a user feels before streamed text starts to appear.