Tokens & Costs
Prompt caching calculator for Claude, GPT and Gemini
Work out what prompt caching really saves on Claude, GPT and Gemini: the hit rate your traffic and cache lifetime will get, what writes and reads cost, and when it pays off. Free, and it runs in your browser.
$2 input · $0.10 cached · $2.50 cache write · $10 output per 1M. 209 models have a cached-input price.
Identical at the start of every request: tools, system prompt, fixed documents
The question and anything that changes
Including any thinking tokens
24 for round-the-clock traffic
Writes cost $2.50 per 1M (1.25× input); each hit refreshes the lifetime.
- Hit rate (estimated)
- 99.9%
- Cache writes a day
- 1.0
- Cache reads a day
- 2,999
- Requests a minute while active
- 4.17
- Per day with / without
- $16.82 / $73.80
A cache write pays for itself after 1 read, which means a hit rate above 20.8%.
With this lifetime (5 minutes) and 12 hours of traffic a day, caching starts saving at about 35 requests a day.
How traffic changes the cost of Claude Sonnet 5.5
Your workload: 3K requests a day, per 1,000 requests$5.61with caching$24.60without
Cheapest with caching for this workload
| # | Model | Lifetime | Hit rate | Per month | Action |
|---|---|---|---|---|---|
| 1 | GPT-5 NanoOpenAI | In-memory (about 5 minutes idle) | 99.9% | $20.54 | |
| 2 | Gemini 2.5 Flash LiteGoogle | Assume 5 minutes | 99.9% | $26.49 | |
| 3 | GPT-6 LunaOpenAI | 30 minutes | 99.9% | $30.15 | |
| 4 | GPT-6 Luna ProOpenAI | 30 minutes | 99.9% | $30.15 | |
| 5 | Claude Haiku 5.5Anthropic | 5 minutes | 99.9% | $30.15 | |
| 6 | GPT-4.1 NanoOpenAI | In-memory (about 5 minutes idle) | 99.9% | $40.17 | |
| 7 | GPT-5.6 LunaOpenAI | 30 minutes | 99.9% | $67.59 | |
| 8 | GPT-5.6 Luna ProOpenAI | 30 minutes | 99.9% | $67.59 | |
| 9 | GPT-5.4 NanoOpenAI | In-memory (about 5 minutes idle) | 99.9% | $69.40 | |
| 10 | Gemini 3.1 Flash LiteGoogle | Assume 5 minutes | 99.9% | $84.47 |
Each model is priced at its cheapest published lifetime for your traffic; where the provider publishes none (Gemini’s implicit cache, unchecked providers), at the assumed lifetime: 5 minutes unless you chose another for such a model above. Cheapest isn’t best: quality and speed differ a lot, so test candidates on your own prompts.
Steps
How to use the prompt caching calculator
- Pick the model. Only models with a published cached-input price are listed.
- Enter the cacheable prefix (tools, system prompt, fixed documents), the new input per request and the output, including thinking tokens.
- Set requests per day and how many hours a day the traffic runs, so the request rate per minute is realistic.
- Choose the cache lifetime where the provider offers a choice, or set the hit rate by hand if you have measured it.
- Read the monthly cost with and without caching, check the chart and the cheapest caching models, and copy the link to share it.
Method
How it works
Prompt caching bills the repeated start of a prompt, the prefix, at a reduced price when it was seen recently. Most calculators assume a hit rate and stop there. This one works out the hit rate from your traffic and the cache lifetime, then applies each provider’s write and read prices from our daily price data (209 models; 74 from Anthropic, OpenAI and Google, whose caching rules we checked on 2026-10-11).
The cost formula
With N requests a day, a prefix of P tokens, Q new input tokens and O output tokens per request:
without = N × (P + Q) × input + N × O × outputwith = writes × P × write + reads × P × cached + N × Q × input + N × O × output
Prices are per million tokens. A miss writes the prefix: on Anthropic at 1.25× input for the 5-minute cache or 2× for the 1-hour cache; on OpenAI GPT-5.6 and later at 1.25×; on earlier OpenAI models and Gemini’s implicit cache at the normal input price. Monthly figures use 365 ÷ 12 days. Long-context price tiers apply to the whole request, and the cached and write prices scale with them, as Anthropic’s Haiku 5.5 price table shows.
How the hit rate is estimated
Every request, hit or miss, leaves the prefix in the cache for another T minutes (the lifetime), because Anthropic and OpenAI refresh an entry each time it’s used (for Gemini and other providers, we assume the same). So a request finds the cache warm when the previous request arrived less than T minutes earlier. If requests arrive at random at λ per minute (requests per day ÷ active minutes), the gap between them is exponentially distributed and the chance it is shorter than T is 1 − e^(−λT). The first request of each day always writes, unless the overnight gap is shorter than the lifetime. For example, one request every 10 minutes against a 5-minute cache gives λT = 0.5 and a hit rate of only 39%. Our unit tests check the formula against a simulation of 200,000 random requests.
Where the estimate can be wrong
- Bursty or regular traffic. Clustered requests hit more often than random ones with the same daily total; evenly spaced requests are all-or-nothing (every gap under T hits, every gap over T misses).
- Concurrency. Anthropic notes that an entry only becomes available after the first response begins, so requests fired in parallel before then each write.
- Provider-side behaviour. OpenAI says that on models before GPT-5.6, traffic above about 15 requests a minute can be routed to machines without your cache, and that nothing guarantees a hit. Google says Gemini’s implicit caching has no cost-saving guarantee and doesn’t publish its lifetime, so the lifetime for Gemini is an assumption you choose.
- Only the fixed prefix is counted. In a chat, the growing history can be cached too; for that, measure your cached share and enter it in the LLM cost calculator.
If you have measured your hit rate from the API’s cached-token fields, set it by hand and the traffic model is skipped. Not included: Gemini’s explicit caches (billed for storage per token-hour), batch discounts, regional price multipliers and rate-limit effects.
Break-even
If a write costs W times the input price and a read R times, a write pays for itself after more than (W − 1) ÷ (1 − R) reads, and caching saves money once the hit rate passes (W − 1) ÷ (W − R). With no write premium (W = 1), any hit is a saving. The prompt caching guide explains these rules in more depth, and the token counter tells you how long your prefix is.
Examples
Worked examples
Support bot: a 10,000-token system prompt, 3,000 questions a day over 12 hours
The prefix holds the instructions and tool definitions; each question adds 300 tokens and the reply is 400. At about 4 requests a minute the cache never goes cold, so nearly every request is a read and only the first one each morning writes. The saving is limited by what caching can’t touch: output and new input.
| Model | Lifetime | Hit rate | Writes a day | Without caching | With caching | Saving |
|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 | 5 minutes | 99.9% | 1.0 | $2,245 | $511.73 | 77% |
| GPT-6.1 Sol | 30 minutes | 99.9% | 1.0 | $2,245 | $511.73 | 77% |
| Gemini 3.8 Flash | Assume 5 minutes | 99.9% | 1.0 | $841.78 | $226.05 | 73% |
RAG over a fixed document set: 30,000 tokens of documents, 400 questions a day over 10 hours
When the same documents sit at the top of every prompt, caching pays well. At 0.67 requests a minute, a 5-minute cache still expires now and then, which is where a longer lifetime earns its higher write price. This only works if retrieval returns the same set in the same order; per-question chunks belong after the fixed prefix.
| Model | Lifetime | Hit rate | Writes a day | Without caching | With caching | Saving |
|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 | 5 minutes | 96.1% | 15 | $815.17 | $155.03 | 81% |
| Claude Sonnet 5.5 | 1 hour | 99.7% | 1.0 | $815.17 | $125.23 | 85% |
| GPT-6.1 Sol | 30 minutes | 99.7% | 1.0 | $815.17 | $123.86 | 85% |
| Gemini 3.8 Flash | Assume 5 minutes | 96.1% | 15 | $305.69 | $68.70 | 78% |
Low-traffic internal tool: a 8,000-token prompt, 20 uses a day over 8 hours
One use every 24 minutes on average is too slow for a 5-minute cache: most requests find it expired and pay the write premium again. On Claude Sonnet 5.5 the 5-minute cache hits 18% of the time and costs 3% more than no caching; the 1-hour cache hits 87% and saves 52%. At this volume the bill is small either way, so caching barely matters in dollars.
| Model | Lifetime | Hit rate | Writes a day | Without caching | With caching | Saving |
|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 | 5 minutes | 17.8% | 16 | $13.14 | $13.49 | -3% |
| Claude Sonnet 5.5 | 1 hour | 87.2% | 2.6 | $13.14 | $6.32 | 52% |
| GPT-6.1 Sol | 30 minutes | 67.7% | 6.4 | $13.14 | $7.66 | 42% |
| Gemini 3.8 Flash | Assume 5 minutes | 17.8% | 16 | $4.93 | $4.34 | 12% |
Per month, prices as of 2026-10-11; provider rules checked 2026-10-11. Gemini rows assume a 5-minute implicit-cache lifetime, which Google doesn’t publish. The cost estimation guide walks through the same maths for a whole conversation, and reducing Claude Code token usage shows caching in a coding agent.
How each provider bills prompt caching
| Rule | Anthropic (Claude) | OpenAI GPT-5.6 and later | Earlier OpenAI models | Google Gemini (implicit) |
|---|---|---|---|---|
| Turned on by | You: cache_control, automatic or up to 4 breakpoints | Default; implicit or explicit breakpoints | Default | Default on Gemini 2.5 and newer |
| Lifetime | 5 minutes or 1 hour, refreshed on each use | 30 minutes after the last write or reuse | About 5 to 10 minutes idle; about 30 minutes with 24h retention | Not published |
| Cache write | 1.25× input (5 min), 2× (1 hour) | 1.25× input | No charge (normal input) | No charge (normal input) |
| Cache read | 0.1× input; 0.05× Opus 5.5 and Sonnet 5.5; 0.025× Fable 5.1 | 0.1× input; 0.05× GPT-6.1 Sol | Per model | Per model; no cost-saving guarantee |
| Minimum prefix | 512 tokens on the newest models, up to 4,096 on older ones | 1,024 tokens | Varies by request settings | 4,096 (3.8 Flash, 3.6 Flash, 3.1 Pro Preview); 2,048 (2.5) |
From each provider’s caching and pricing documentation, checked 2026-10-11. Read prices per model come from our data; for example Claude Sonnet 5.5 reads at $0.10 per 1M against $2.
FAQ
Frequently asked questions
How much does prompt caching save?
On the cached prefix, reads cost 90% less than input on most Claude and GPT-5.6+ models and 97.5% less on Claude Fable 5.1, but only 50% less on older models such as GPT-4o. The whole bill falls less, because new input and output are billed normally: our support-bot example on Claude Sonnet 5.5 saves 77%. Slow traffic can save nothing.
How long does a prompt cache last?
Anthropic: 5 minutes by default, or 1 hour for a higher write price, refreshed each time the entry is used. OpenAI: 30 minutes after the last write or reuse on GPT-5.6 and later; on earlier models typically 5 to 10 minutes idle in memory, or about 30 minutes with 24-hour retention. Google doesn’t publish a lifetime for Gemini’s implicit cache.
Does OpenAI charge for cache writes?
On GPT-5.6 and later models, yes: OpenAI bills a cache write at 1.25× the normal input price, and reads at 0.1× (0.05× on GPT-6.1 Sol). Earlier models have no write charge: a miss costs the normal input price and a hit costs the model’s cached-input price, so caching on those models never costs more than not caching.
Is Claude’s 1-hour cache worth the higher price?
It costs 2× the input price to write instead of 1.25×, so it needs two cache reads to pay off instead of one. It wins when the gap between requests is often longer than 5 minutes but shorter than an hour, such as an internal tool used a few times an hour. With steady traffic, the 5-minute cache stays warm on its own and is cheaper.
Why is my real cache hit rate lower than this estimate?
Usually the prefix isn’t identical: a timestamp, user name or reordered JSON near the top breaks every hit. Other causes: the prefix is below the model’s minimum, parallel requests sent before the first response started, traffic routed to machines that don’t hold your cache, or bursts followed by long idle gaps. Check the cached-token fields in each API response.
Does prompt caching make responses faster?
Usually, for long prompts. Anthropic says you will generally see a faster time to first token for long documents, and that its 5-minute and 1-hour caches behave the same for latency. OpenAI says caching reduces the time spent processing input before the response starts. Google publishes no latency figure for Gemini. Writing the reply takes as long as before, so the gain is largest for a long prompt with a short answer.
What is the minimum prompt length for caching?
Anthropic caches prefixes of 512 tokens on its newest models (Fable 5.1, Opus 5.5, Sonnet 5.5, Haiku 5.5) and 1,024 to 4,096 on older ones. OpenAI’s minimum is 1,024 tokens on GPT-5.6 and later. Gemini’s implicit cache needs 4,096 tokens on Gemini 3.8 Flash, 3.6 Flash and 3.1 Pro Preview, and 2,048 on Gemini 2.5. Shorter prefixes are simply not cached.
Related
Related tools
- LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- AI Model ComparisonCompare prices, context windows and features across models.
- Context Window CheckerSee whether your text fits each model’s context window.