Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

Tokens & Costs

Prompt caching calculator for Claude, GPT and Gemini

Work out what prompt caching really saves on Claude, GPT and Gemini: the hit rate your traffic and cache lifetime will get, what writes and reads cost, and when it pays off. Free, and it runs in your browser.

$2 input · $0.10 cached · $2.50 cache write · $10 output per 1M. 209 models have a cached-input price.

Start from an example

Identical at the start of every request: tools, system prompt, fixed documents

The question and anything that changes

Including any thinking tokens

24 for round-the-clock traffic

Writes cost $2.50 per 1M (1.25× input); each hit refreshes the lifetime.

Hit rate
Per month with caching$511.73
Without caching$2,245
Saving per month$1,733
Saving77.2%
Hit rate (estimated)
99.9%
Cache writes a day
1.0
Cache reads a day
2,999
Requests a minute while active
4.17
Per day with / without
$16.82 / $73.80

A cache write pays for itself after 1 read, which means a hit rate above 20.8%.

With this lifetime (5 minutes) and 12 hours of traffic a day, caching starts saving at about 35 requests a day.

Prices 2026-10-11 · rules 2026-10-11

How traffic changes the cost of Claude Sonnet 5.5

With cachingWithout caching

Your workload: 3K requests a day, per 1,000 requests$5.61with caching$24.60without

Cost per 1,000 requests (y) against requests per day on a log scale (x), with your other inputs fixed. The dot is your workload. More traffic keeps the cache warm, so the cached line falls; the uncached line is flat.

Cheapest with caching for this workload

ModelPer monthAction
GPT-5 NanoOpenAI$20.54
Gemini 2.5 Flash LiteGoogle$26.49
GPT-6 LunaOpenAI$30.15
GPT-6 Luna ProOpenAI$30.15
Claude Haiku 5.5Anthropic$30.15
GPT-4.1 NanoOpenAI$40.17
GPT-5.6 LunaOpenAI$67.59
GPT-5.6 Luna ProOpenAI$67.59
GPT-5.4 NanoOpenAI$69.40
Gemini 3.1 Flash LiteGoogle$84.47

Each model is priced at its cheapest published lifetime for your traffic; where the provider publishes none (Gemini’s implicit cache, unchecked providers), at the assumed lifetime: 5 minutes unless you chose another for such a model above. Cheapest isn’t best: quality and speed differ a lot, so test candidates on your own prompts.

Steps

How to use the prompt caching calculator

  1. Pick the model. Only models with a published cached-input price are listed.
  2. Enter the cacheable prefix (tools, system prompt, fixed documents), the new input per request and the output, including thinking tokens.
  3. Set requests per day and how many hours a day the traffic runs, so the request rate per minute is realistic.
  4. Choose the cache lifetime where the provider offers a choice, or set the hit rate by hand if you have measured it.
  5. Read the monthly cost with and without caching, check the chart and the cheapest caching models, and copy the link to share it.

Method

How it works

Prompt caching bills the repeated start of a prompt, the prefix, at a reduced price when it was seen recently. Most calculators assume a hit rate and stop there. This one works out the hit rate from your traffic and the cache lifetime, then applies each provider’s write and read prices from our daily price data (209 models; 74 from Anthropic, OpenAI and Google, whose caching rules we checked on 2026-10-11).

The cost formula

With N requests a day, a prefix of P tokens, Q new input tokens and O output tokens per request:

without = N × (P + Q) × input + N × O × output
with = writes × P × write + reads × P × cached + N × Q × input + N × O × output

Prices are per million tokens. A miss writes the prefix: on Anthropic at 1.25× input for the 5-minute cache or 2× for the 1-hour cache; on OpenAI GPT-5.6 and later at 1.25×; on earlier OpenAI models and Gemini’s implicit cache at the normal input price. Monthly figures use 365 ÷ 12 days. Long-context price tiers apply to the whole request, and the cached and write prices scale with them, as Anthropic’s Haiku 5.5 price table shows.

How the hit rate is estimated

Every request, hit or miss, leaves the prefix in the cache for another T minutes (the lifetime), because Anthropic and OpenAI refresh an entry each time it’s used (for Gemini and other providers, we assume the same). So a request finds the cache warm when the previous request arrived less than T minutes earlier. If requests arrive at random at λ per minute (requests per day ÷ active minutes), the gap between them is exponentially distributed and the chance it is shorter than T is 1 − e^(−λT). The first request of each day always writes, unless the overnight gap is shorter than the lifetime. For example, one request every 10 minutes against a 5-minute cache gives λT = 0.5 and a hit rate of only 39%. Our unit tests check the formula against a simulation of 200,000 random requests.

Where the estimate can be wrong

  • Bursty or regular traffic. Clustered requests hit more often than random ones with the same daily total; evenly spaced requests are all-or-nothing (every gap under T hits, every gap over T misses).
  • Concurrency. Anthropic notes that an entry only becomes available after the first response begins, so requests fired in parallel before then each write.
  • Provider-side behaviour. OpenAI says that on models before GPT-5.6, traffic above about 15 requests a minute can be routed to machines without your cache, and that nothing guarantees a hit. Google says Gemini’s implicit caching has no cost-saving guarantee and doesn’t publish its lifetime, so the lifetime for Gemini is an assumption you choose.
  • Only the fixed prefix is counted. In a chat, the growing history can be cached too; for that, measure your cached share and enter it in the LLM cost calculator.

If you have measured your hit rate from the API’s cached-token fields, set it by hand and the traffic model is skipped. Not included: Gemini’s explicit caches (billed for storage per token-hour), batch discounts, regional price multipliers and rate-limit effects.

Break-even

If a write costs W times the input price and a read R times, a write pays for itself after more than (W − 1) ÷ (1 − R) reads, and caching saves money once the hit rate passes (W − 1) ÷ (W − R). With no write premium (W = 1), any hit is a saving. The prompt caching guide explains these rules in more depth, and the token counter tells you how long your prefix is.

Examples

Worked examples

Support bot: a 10,000-token system prompt, 3,000 questions a day over 12 hours

The prefix holds the instructions and tool definitions; each question adds 300 tokens and the reply is 400. At about 4 requests a minute the cache never goes cold, so nearly every request is a read and only the first one each morning writes. The saving is limited by what caching can’t touch: output and new input.

Support bot example
ModelLifetimeHit rateWrites a dayWithout cachingWith cachingSaving
Claude Sonnet 5.55 minutes99.9%1.0$2,245$511.7377%
GPT-6.1 Sol30 minutes99.9%1.0$2,245$511.7377%
Gemini 3.8 FlashAssume 5 minutes99.9%1.0$841.78$226.0573%

RAG over a fixed document set: 30,000 tokens of documents, 400 questions a day over 10 hours

When the same documents sit at the top of every prompt, caching pays well. At 0.67 requests a minute, a 5-minute cache still expires now and then, which is where a longer lifetime earns its higher write price. This only works if retrieval returns the same set in the same order; per-question chunks belong after the fixed prefix.

RAG example
ModelLifetimeHit rateWrites a dayWithout cachingWith cachingSaving
Claude Sonnet 5.55 minutes96.1%15$815.17$155.0381%
Claude Sonnet 5.51 hour99.7%1.0$815.17$125.2385%
GPT-6.1 Sol30 minutes99.7%1.0$815.17$123.8685%
Gemini 3.8 FlashAssume 5 minutes96.1%15$305.69$68.7078%

Low-traffic internal tool: a 8,000-token prompt, 20 uses a day over 8 hours

One use every 24 minutes on average is too slow for a 5-minute cache: most requests find it expired and pay the write premium again. On Claude Sonnet 5.5 the 5-minute cache hits 18% of the time and costs 3% more than no caching; the 1-hour cache hits 87% and saves 52%. At this volume the bill is small either way, so caching barely matters in dollars.

Internal tool example
ModelLifetimeHit rateWrites a dayWithout cachingWith cachingSaving
Claude Sonnet 5.55 minutes17.8%16$13.14$13.49-3%
Claude Sonnet 5.51 hour87.2%2.6$13.14$6.3252%
GPT-6.1 Sol30 minutes67.7%6.4$13.14$7.6642%
Gemini 3.8 FlashAssume 5 minutes17.8%16$4.93$4.3412%

Per month, prices as of 2026-10-11; provider rules checked 2026-10-11. Gemini rows assume a 5-minute implicit-cache lifetime, which Google doesn’t publish. The cost estimation guide walks through the same maths for a whole conversation, and reducing Claude Code token usage shows caching in a coding agent.

How each provider bills prompt caching

Prompt caching rules by provider
RuleAnthropic (Claude)OpenAI GPT-5.6 and laterEarlier OpenAI modelsGoogle Gemini (implicit)
Turned on byYou: cache_control, automatic or up to 4 breakpointsDefault; implicit or explicit breakpointsDefaultDefault on Gemini 2.5 and newer
Lifetime5 minutes or 1 hour, refreshed on each use30 minutes after the last write or reuseAbout 5 to 10 minutes idle; about 30 minutes with 24h retentionNot published
Cache write1.25× input (5 min), 2× (1 hour)1.25× inputNo charge (normal input)No charge (normal input)
Cache read0.1× input; 0.05× Opus 5.5 and Sonnet 5.5; 0.025× Fable 5.10.1× input; 0.05× GPT-6.1 SolPer modelPer model; no cost-saving guarantee
Minimum prefix512 tokens on the newest models, up to 4,096 on older ones1,024 tokensVaries by request settings4,096 (3.8 Flash, 3.6 Flash, 3.1 Pro Preview); 2,048 (2.5)

From each provider’s caching and pricing documentation, checked 2026-10-11. Read prices per model come from our data; for example Claude Sonnet 5.5 reads at $0.10 per 1M against $2.

FAQ

Frequently asked questions

How much does prompt caching save?

On the cached prefix, reads cost 90% less than input on most Claude and GPT-5.6+ models and 97.5% less on Claude Fable 5.1, but only 50% less on older models such as GPT-4o. The whole bill falls less, because new input and output are billed normally: our support-bot example on Claude Sonnet 5.5 saves 77%. Slow traffic can save nothing.

How long does a prompt cache last?

Anthropic: 5 minutes by default, or 1 hour for a higher write price, refreshed each time the entry is used. OpenAI: 30 minutes after the last write or reuse on GPT-5.6 and later; on earlier models typically 5 to 10 minutes idle in memory, or about 30 minutes with 24-hour retention. Google doesn’t publish a lifetime for Gemini’s implicit cache.

Does OpenAI charge for cache writes?

On GPT-5.6 and later models, yes: OpenAI bills a cache write at 1.25× the normal input price, and reads at 0.1× (0.05× on GPT-6.1 Sol). Earlier models have no write charge: a miss costs the normal input price and a hit costs the model’s cached-input price, so caching on those models never costs more than not caching.

Is Claude’s 1-hour cache worth the higher price?

It costs 2× the input price to write instead of 1.25×, so it needs two cache reads to pay off instead of one. It wins when the gap between requests is often longer than 5 minutes but shorter than an hour, such as an internal tool used a few times an hour. With steady traffic, the 5-minute cache stays warm on its own and is cheaper.

Why is my real cache hit rate lower than this estimate?

Usually the prefix isn’t identical: a timestamp, user name or reordered JSON near the top breaks every hit. Other causes: the prefix is below the model’s minimum, parallel requests sent before the first response started, traffic routed to machines that don’t hold your cache, or bursts followed by long idle gaps. Check the cached-token fields in each API response.

Does prompt caching make responses faster?

Usually, for long prompts. Anthropic says you will generally see a faster time to first token for long documents, and that its 5-minute and 1-hour caches behave the same for latency. OpenAI says caching reduces the time spent processing input before the response starts. Google publishes no latency figure for Gemini. Writing the reply takes as long as before, so the gain is largest for a long prompt with a short answer.

What is the minimum prompt length for caching?

Anthropic caches prefixes of 512 tokens on its newest models (Fable 5.1, Opus 5.5, Sonnet 5.5, Haiku 5.5) and 1,024 to 4,096 on older ones. OpenAI’s minimum is 1,024 tokens on GPT-5.6 and later. Gemini’s implicit cache needs 4,096 tokens on Gemini 3.8 Flash, 3.6 Flash and 3.1 Pro Preview, and 2,048 on Gemini 2.5. Shorter prefixes are simply not cached.