Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Guide · Tokens & costs

Prompt caching explained: how it works on OpenAI, Anthropic and Gemini, and when it pays

Prompt caching lets an LLM provider reuse the work it already did on the start of your prompt. When a request begins with exactly the same tokens as a recent one, that prefix is billed at a fraction of the normal input price and processed faster. It pays off when a long, fixed prefix is reused within minutes.

By Tahir NazirUpdated 11 min read

On this page
  1. What is prompt caching?
  2. How much does prompt caching save?
  3. How prompt caching works on Anthropic, OpenAI and Google
  4. When does prompt caching pay off?
  5. How to structure prompts for cache hits
  6. When prompt caching doesn’t help
  7. How to check that caching is working
  8. Questions people ask

What is prompt caching?

Prompt caching is a provider-side discount for repeated input. Before a model writes anything, it processes your whole prompt into internal state (the attention “keys and values” for every token). If the next request starts with exactly the same tokens, the provider can reuse that saved state for the shared start, the prefix, instead of computing it again. You pay a reduced price for those cached tokens, and the response starts sooner.

Three details matter in practice:

  • It matches prefixes, not whole prompts. Everything up to the first token that differs can be reused; everything after it is processed normally. One changed character near the top of a prompt makes the rest of it uncacheable.
  • It doesn’t change the answer. OpenAI’s documentation is direct: “Prompt caching does not change how the model generates output tokens.” Output is billed at the normal price.
  • It isn’t response caching. The provider doesn’t store or replay answers. Caching whole responses for repeated questions (sometimes called semantic caching) is something you build in your own app.

Caches are private. Anthropic says caches are never shared across organisations and are isolated per workspace on its own API; OpenAI says caches are not shared across organisations.

How much does prompt caching save?

It depends on how much of each prompt is a repeated prefix and how often it repeats. Take a document assistant on Claude Sonnet 5.5: a 10,000-token prefix (instructions plus a document) followed by a 200-token question, with a 300-token answer. The first request writes the prefix to the cache; the next five read it:

The first request pays a write premium; every later one reads the prefix at the cached rate. Over six requests caching saves 64%. Prices as of 2026-10-09.

The write costs $2.50 per million tokens instead of $2, so request 1 is dearer than it would be without caching. From request 2 the prefix costs $0.10 per million, and what remains is mostly the answer: output at $10 per million is now the largest part of each request. That is the general pattern. Caching can remove most of the cost of repeated input, but it does nothing for output. The cost estimation guide applies the same maths to a whole 10-turn conversation.

How prompt caching works on Anthropic, OpenAI and Google

All three providers cache prefixes, but they differ in whether you have to ask for it, how long entries live and how writes are charged. As documented on 8 October 2026:

Prompt caching by provider
Anthropic (Claude)OpenAIGoogle (Gemini)
How it’s turned onOpt-in: a top-level cache_control field (automatic) or up to 4 breakpoints on content blocksOn by default; on GPT-5.6 and later, implicit or explicit breakpoint modesImplicit caching on by default (Gemini 2.5 and newer); explicit caches you create
Minimum prefix512 tokens on Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 5.5; up to 4,096 on older models1,024 tokens on GPT-5.6 and later; varies on earlier modelsImplicit: 4,096 tokens on Gemini 3.5 to 3.8 Flash; 2,048 on Gemini 2.5
How long it lasts5 minutes, refreshed free on each hit; 1-hour optionAt least 30 minutes on GPT-5.6 and later; about 5 to 10 minutes idle on earlier models, with 24-hour retention on someExplicit: a TTL you choose, 1 hour by default
Write cost1.25× input (5 minutes) or 2× (1 hour)1.25× input on GPT-5.6 and later; none on earlier modelsImplicit: none. Explicit: storage billed per token-hour
Read price0.1× input on most models (0.05× on Opus 5.5 and Sonnet 5.5, 0.025× on Fable 5.1); Sonnet 5.5: $0.10 vs $20.1× on most GPT-5.6+ models; GPT-6.1 Sol: $0.10 vs $2Per model; Gemini 3.8 Flash: $0.075 vs $0.75

Rules from each provider’s caching and pricing docs, checked 2026-10-08. Example prices per million tokens from our data as of 2026-10-09.

To compare cached prices across every model, filter the model comparison to models with prompt caching and sort by the cached price column.

Anthropic: you choose what to cache

Claude caches nothing unless you ask. The simplest way is a top-level cache_control: {"type": "ephemeral"}, which places the breakpoint on the last cacheable block and moves it forward as a conversation grows. For finer control you mark up to four blocks yourself. The prefix is built in a fixed order, tools then system then messages, and a change at one level invalidates that level and everything after it. Prompts below the model’s minimum are silently not cached, so check the usage fields. Anthropic also notes that cache hits are not deducted from your rate limit.

OpenAI: automatic, with write charges on the newest models

OpenAI caches automatically. On GPT-5.6 and later, implicit mode places a breakpoint at the end of the latest eligible message, explicit mode lets you mark breakpoints, and entries last at least 30 minutes after the latest write or reuse. Those models charge 1.25× the input price to write. Earlier models have no write charge and keep entries for about 5 to 10 minutes of inactivity, with an optional 24-hour retention on some. Unlike Anthropic, OpenAI counts cached tokens toward your tokens-per-minute limit. Cached state lives on individual machines, and OpenAI notes that traffic above 15 requests per minute can be routed to overflow machines. On models before GPT-5.6, a stable prompt_cache_key helps send related requests to the same cache; from GPT-5.6, routing is automatic.

Google: implicit by default, explicit for guaranteed savings

Gemini 2.5 and newer cache implicitly with nothing to set up, but Google describes implicit caching as having no cost-saving guarantee. Explicit caching lets you create a cache object (a document, a video, a long system instruction), refer to it in later requests and get a guaranteed discount. You pay the reduced rate for cached tokens each time they’re used, plus storage for as long as the cache lives. Google’s tips: put large, common content at the start of the prompt and send requests with a similar prefix close together.

When does prompt caching pay off?

Caching pays off once the savings on reads cover the extra cost of the write. If a write costs W times the input price and a read costs R times, a prefix needs more than (W − 1) ÷ (1 − R) reads within its lifetime to come out ahead. On Claude Sonnet 5.5, W is 1.25 and R is 0.05:

  • 5-minute cache: (1.25 − 1) ÷ (1 − 0.05) = 0.26, so a single read pays for the write.
  • 1-hour cache: with Anthropic’s 2× write price, (2 − 1) ÷ (1 − 0.05) = 1.05, so it takes 2 reads. Anthropic’s pricing page reaches the same answer for its standard 0.1× read price: caching “pays off after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write)”.

The lifetime is the catch. Suppose the same 10,000-token document gets one question every 10 minutes. A 5-minute cache has always expired, so every request pays the write premium: $0.15 an hour for the prefix, against $0.12 with no caching at all. A 1-hour cache writes once and reads five times, which brings the hour to $0.045. Match the cache lifetime to the gap between your requests, not to the length of a session.

Google’s explicit caches follow different maths: storage is charged per token per hour, so the cache pays off when the hits in an hour save more than an hour of storage. Divide the storage price by the per-token saving on each hit to get the hits per hour you need, using the prices on Google’s pricing page for your model.

Free toolLLM API cost calculatorSet the cached share of your prompt to see what caching does to your monthly bill, with the cached price for each model.

How to structure prompts for cache hits

  1. Put stable content first. Tool definitions, the system prompt, long documents and few-shot examples go at the top; the user’s question and anything per-request go last.
  2. Keep the prefix byte-for-byte identical. A timestamp, request ID or user name near the top breaks every hit. So does serialising the same JSON with keys in a different order.
  3. Don’t change settings mid-conversation. On Claude, changing tool definitions invalidates every cache level, and changing thinking or effort settings typically invalidates the cached messages.
  4. Warm the cache before a burst. Anthropic notes that a cache entry only becomes available after the first response begins, so send one request and wait for it before firing parallel ones.
  5. Append, don’t rewrite. In chats and agents, add new turns at the end. Summarising or trimming old history changes the prefix and forces a fresh write.

Counting the stable part of your prompt with the token counter tells you whether it clears the provider’s minimum and what share of each request it is.

When prompt caching doesn’t help

  • Short prompts. Below the minimum (512 to 4,096 tokens depending on provider and model) nothing is cached.
  • Slow or bursty-then-idle traffic. If the prefix is reused less often than the cache lives, you pay write premiums for nothing.
  • Prompts that differ early. Per-user data or retrieved documents at the top leave little shared prefix. Move them below the fixed instructions.
  • Output-heavy work. When most of the cost is the answer (long drafts, code generation, heavy reasoning), caching the input moves the total only a little. A model with cheaper output does more; see the cheapest LLM APIs.
  • Batch jobs, unless prepared. Anthropic says cache hits in batches are best-effort because requests run concurrently in any order; it suggests writing the prefix to a 1-hour cache first.

How to check that caching is working

Don’t assume hits; read them from the response. Anthropic returns cache_creation_input_tokens (written) and cache_read_input_tokens (read) alongside input_tokens, which counts only the tokens after the last breakpoint. OpenAI reports cached tokens under usage.input_tokens_details.cached_tokens, inside the input_tokens total. Gemini reports cached tokens in the response’s usage_metadata. Pricing a Claude response from its usage:

Python: cost of a Claude response from its usage fields
def claude_cost(usage, input_price, output_price, write=1.25, read=0.1):
    """Prices in USD per 1M tokens. write/read: the model's cache multipliers."""
    return (
        usage["input_tokens"] * input_price
        + usage["cache_creation_input_tokens"] * input_price * write
        + usage["cache_read_input_tokens"] * input_price * read
        + usage["output_tokens"] * output_price
    ) / 1_000_000

# A cache hit: 10,000 prefix tokens read, a 200-token question, a 300-token answer.
usage = {"input_tokens": 200, "cache_creation_input_tokens": 0,
         "cache_read_input_tokens": 10_000, "output_tokens": 300}

# Example prices (not a real model): $1 in, $5 out per 1M tokens.
hit = claude_cost(usage, 1.00, 5.00)
miss = (10_200 * 1.00 + 300 * 5.00) / 1_000_000
total_in = sum(usage[k] for k in ("input_tokens", "cache_creation_input_tokens", "cache_read_input_tokens"))
print(f"with cache ${hit:.4f}, without ${miss:.4f}, "
      f"cached share {usage['cache_read_input_tokens'] / total_in:.0%}")

# Output: with cache $0.0027, without $0.0117, cached share 98%

Track the cached share over a day of traffic. If it’s far below the share of your prompt that is meant to be stable, something in the prefix is changing between requests. Once you know your real cached share, enter it in the LLM cost calculator to see the monthly effect.

FAQ

Questions people ask

Does prompt caching change the model’s answers?

No. Caching reuses the provider’s internal processing of an identical prefix, so the model sees exactly the same input. OpenAI’s documentation states that prompt caching does not change how the model generates output tokens. Only the input price and the time to the first token change.

How long does a prompt cache last?

It depends on the provider. Anthropic’s default is 5 minutes, refreshed free on every hit, with a 1-hour option at a higher write price. OpenAI keeps entries at least 30 minutes on GPT-5.6 and later, and about 5 to 10 minutes of inactivity on earlier models. Gemini’s explicit caches last for the TTL you set, 1 hour by default.

Is OpenAI prompt caching automatic?

Yes. Prompt caching is on by default for supported OpenAI models, with no code changes needed for prompts above the minimum length (1,024 tokens on GPT-5.6 and later). On GPT-5.6 and later you can also choose explicit breakpoints, and those models charge 1.25 times the input price for cache writes.

What is the minimum prompt length for caching?

Anthropic’s minimum is 512 tokens on Claude Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 5.5, and up to 4,096 on older Claude models. OpenAI’s is 1,024 tokens on GPT-5.6 and later. Gemini’s implicit cache needs 4,096 tokens on Gemini 3.5 to 3.8 Flash and 2,048 on Gemini 2.5. Shorter prompts simply aren’t cached.

Can other customers see my cached prompts?

No. Cache entries are keyed to the exact prompt and kept within your account. Anthropic says caches are never shared across organisations and are isolated per workspace on its own API, and OpenAI says caches are not shared across organisations.

Do cached tokens count toward rate limits?

It differs. Anthropic says cache hits are not deducted against your rate limit, and lists better rate-limit use as one reason to choose its 1-hour cache. OpenAI’s documentation says cached input tokens still count toward your tokens-per-minute limit. Check your provider’s rate-limit docs before planning capacity around caching.

Try it

Tools from this guide

Keep reading