Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Tokens and cost

What is prompt caching?

Also called: context caching

Definition

Prompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.

Explained

How it works

Most requests to a model repeat a long opening: the system prompt, tool definitions, a document, or the conversation so far. With caching on, the provider keeps the model’s work on that opening (the prefix) for a few minutes. If the next request starts with the identical tokens, it reads the prefix from the cache instead of processing it again.

Only an exact prefix counts. One changed character early in the prompt, such as a timestamp in the system prompt, means everything after it is new. OpenAI and Gemini cache automatically above a minimum length; Anthropic caches what you mark with cache_control, or the last cacheable block when you turn on automatic caching. Caches expire after a short idle time (Anthropic’s default is 5 minutes, refreshed on every hit), and some providers charge extra for the first write.

Example

A 20,000-token prefix reused 100 times

A support bot sends a 20,000-token system prompt and help-centre extract with every question, then a 300-token question, and gets a 400-token answer. On Claude Sonnet 5.5, input costs $2 per million tokens, a cache write $2.50 and a cache read $0.10.

Without caching, 100 questions cost $4.46. With caching, the first request writes the prefix and the other 99 read it, for $0.71 in total: 84% less, from one setting. The answers cost the same either way, because caching only discounts input.

100 requests on Claude Sonnet 5.5
Without cachingWith caching
Prefix tokens billed at full price2,000,0000
Prefix tokens written to cache020,000
Prefix tokens read from cache01,980,000
Total cost$4.46$0.71

Prices from our daily data, 2026-10-11. Assumes the requests arrive often enough to keep the cache warm.

Cost and quality

Why it matters

Input is usually most of an agent’s or chatbot’s bill, because the same context is sent on every turn. Caching is often the single largest saving available without changing the model, and it also cuts the time to the first token on long prompts.

It only pays when the prefix is reused before it expires. Slow traffic, or a prompt that changes near the top, can mean paying the write premium with no reads to recover it.

Don’t mix up

Common confusions

Prompt caching vs response caching
A response cache in your own code returns a stored answer to an identical question without calling the model. Prompt caching still calls the model and generates a fresh answer; only the processing of the repeated input is reused.
Prompt caching vs the KV cache
The KV cache is the memory a model keeps during a single request. Prompt caching keeps that work between requests, on the provider’s side, and bills it at a discount.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary

Written by Tahir Nazir. Checked .

How this was checked: Cost example computed from our daily price data. Cache rules (5-minute default lifetime, cache_control, automatic caching, OpenAI’s automatic caching above 1,024 tokens on GPT-5.6 and later) checked against Anthropic’s and OpenAI’s documentation on 2026-10-11.