Guide · Tokens & costs
How to estimate LLM API costs: the formula, worked examples and the traps
To estimate an LLM API bill, multiply the input tokens in a request by the input price and the output tokens by the output price (both quoted per million tokens), add the two, then multiply by your requests per month. Most surprises come from tokens you don’t see: re-sent chat history, hidden reasoning and retries.
By Tahir NazirUpdated 10 min read
On this page
How is LLM API cost calculated?
Every major LLM API bills by the token, with one price for input tokens (everything you send) and another for output tokens (everything the model writes). Prices are quoted in US dollars per million tokens, so the cost of one request is:
cost = (uncached input × input price
+ cached input × cached price
+ output × output price) ÷ 1,000,000Input includes far more than the user’s message: the system prompt, tool definitions, retrieved documents and the conversation so far. Output includes any hidden reasoning the model does before it answers. Here is one realistic request on Claude Sonnet 5.5, the model our calculator opens with:
Share of the cost
- New inputuser message + recent history1,000 × $2/M$0.002
- Cached inputsystem prompt read from cache2,000 × $0.10/M$0.0002
- Visible outputthe answer you see400 × $10/M$0.004
- Thinking outputreasoning, billed as output600 × $10/M$0.006
Claude Sonnet 5.5 · list prices as of 2026-10-09
The cached system prompt is the largest part of the request and the smallest part of the bill, because cached input on this model costs $0.10 per million tokens against $2 for fresh input. The thinking tokens cost more than the answer itself. Those two effects, discounts on repeated input and surcharges hidden in output, decide most real bills.
How to estimate your monthly bill, step by step
- Build one real request. Assemble the full prompt your app will send: system prompt, tool definitions, retrieved context, typical history and a typical user message. Count it with the token counter for the model you plan to use, because each model family splits text differently.
- Measure output, don’t guess it. Run 50 to 100 representative prompts and log the output token counts the API returns. Use the average for budgeting and the high end for setting
max_tokens. - Price one request. Apply the formula above with the model’s current prices. Split input into cached and uncached parts if your prompts share a fixed start.
- Multiply by volume. Multiply by requests per day, then by 30.4 (365 ÷ 12) for an average month.
- Add a margin. Retries, unusually long inputs and growing histories push real bills above the average case. Compare your estimate with the provider’s usage dashboard after the first week and correct it.
Once traffic is live, compute cost from the usage numbers in each response instead of from estimates. The providers report them differently: on OpenAI, input_tokens includes the cached tokens shown under input_tokens_details.cached_tokens, while on Anthropic, input_tokens counts only the uncached tail and cache reads and writes are reported separately. A minimal version, which leaves out cache-write surcharges:
def request_cost(usage, price):
"""usage: token counts from the API response. price: USD per 1M tokens."""
uncached = usage["input"] - usage.get("cached", 0)
return (
uncached * price["input"]
+ usage.get("cached", 0) * price["cached"]
+ usage["output"] * price["output"] # output includes reasoning tokens
) / 1_000_000
# Example prices (not a real model): $1 in, $0.10 cached, $5 out per 1M tokens.
price = {"input": 1.00, "cached": 0.10, "output": 5.00}
usage = {"input": 3_000, "cached": 2_000, "output": 1_000}
per_request = request_cost(usage, price)
per_month = per_request * 1_000 * 365 / 12 # 1,000 requests a day
print(f"${per_request:.4f} per request, ${per_month:,.2f} per month")
# Output: $0.0062 per request, $188.58 per monthHow much does an LLM API cost? Four worked examples
The same four workloads that the calculator offers as presets, priced on four current models. Each row is a different shape of traffic, and the shape matters as much as the model:
| Workload (tokens per request) | Claude Opus 5.5 | GPT-6.1 Sol | Gemini 3.8 Flash | DeepSeek V4.1 Flash |
|---|---|---|---|---|
| Support chatbot: 1,500 in, 400 out, 1,000 a day, 50% cached | $339.15 | $169.57 | $64.45 | $21.58 |
| RAG app: 6,000 in, 500 out, 500 a day, 20% cached | $447.73 | $223.87 | $84.63 | $31.13 |
| Coding agent: 40,000 in, 2,000 out, 200 a day, 80% cached | $476.93 | $238.47 | $96.73 | $30.37 |
| Document summariser: 8,000 in, 600 out, 300 a day, batch API | $200.75 | $100.38 | $37.64 | $10.02 |
Computed with the cost calculator’s maths from list prices as of 2026-10-09. “Cached” assumes that share of each prompt is read from a warm cache; months are 30.4 days.
- Support chatbot. Short prompts, a paragraph out, high volume. Half the prompt is a fixed system prompt, which is exactly what caching is for.
- RAG over documents. Retrieval-augmented generation pastes the most relevant chunks of your documents into each prompt, so input dominates. Halving the number of retrieved chunks roughly halves the input cost.
- Coding agent. Very large prompts (files, tool results, history) on every step, mostly cached. Without caching, the same workload costs 2.4 to 2.9 times as much on these four models.
- Batch summarisation. Long documents in, short summaries out, nobody waiting. It runs through the batch API at half price.
Within each row, the dearest model costs 14 to 20 times as much as the cheapest. Choosing the model is the biggest lever you have, which is why it pays to shortlist cheaper models in the model comparison or our daily ranking of the cheapest LLM APIs, then test them on your own prompts.
Why chats and agents cost more than you expect
A model keeps nothing between requests, so every new message in a chat is processed together with the system prompt and the entire conversation so far, and all of it is billed as input. That holds even when the provider stores the conversation for you: OpenAI says that with previous_response_id, all previous input tokens in the chain are billed as input tokens. Input per turn grows with each turn, and the total grows faster than the number of turns.
Turn
- Turn 1 input
- 1,600 tokens
- Turn 10 input
- 5,200 tokens
- Total input, 10 turns
- 34,000 tokens
Output over the same 10 turns: 3,000 tokens.
On Claude Sonnet 5.5 that conversation costs $0.098 without caching. With caching, each turn reads the previous prompt from the cache and pays the cache-write price only for what is new since the last turn, which brings it to about $0.0459. The prompt caching guide explains the write and read prices.
Agents grow faster still. Every tool call adds its arguments and its result (a file, a page of search results, a test log) to the history, and an agent can take dozens of steps for one task. Some models also carry earlier reasoning forward: Anthropic’s documentation says Claude Opus 4.5 and models numbered 4.6 and higher keep prior turns’ thinking blocks in context and bill them as input. To estimate an agent, log the usage of a few complete tasks and price the whole task, not one step.
Are reasoning tokens billed?
Yes. Reasoning (or “thinking”) models work through a problem in tokens before they answer, and all three major providers bill those tokens as output, whether or not you see them:
- OpenAI says reasoning tokens are not visible via the API but “still occupy space in the model’s context window and are billed as output tokens”. They are reported as
output_tokens_details.reasoning_tokens. - Anthropic reports how many billed output tokens were reasoning in
usage.output_tokens_details.thinking_tokens. - Google states on its Gemini pricing page that output prices include thinking tokens.
OpenAI notes that a model may generate “anywhere from a few hundred to tens of thousands of reasoning tokens” depending on the problem. If GPT-6.1 Sol writes a 300-token answer after 3,000 tokens of reasoning, the output costs $0.033 instead of $0.003: 11 times what the visible answer suggests. Measure reasoning on your own prompts, and use the provider’s effort or reasoning settings to keep it in proportion to the task.
Caching, batch and long-context pricing
List prices are only the starting point. Three pricing rules can move a bill by large factors in either direction:
Prompt caching
When requests start with the same text, providers can serve that prefix from a cache. On the popular models in our data, reading cached input costs 2% to 27% of the normal input price. Some providers charge extra to write the cache, and the cache only lasts minutes to hours, so the real saving depends on how often the prefix repeats. See prompt caching explained for the break-even maths.
Batch APIs
OpenAI, Anthropic and Google all offer asynchronous batch processing at 50% off standard prices. OpenAI says each batch completes within 24 hours, often sooner; Anthropic says most batches finish in less than an hour. For the summarisation workload above, Gemini 3.8 Flash costs $37.64 a month through batch against $75.28 without it. Anything nobody is waiting for (nightly reports, backfills, evaluations, classification) belongs in a batch.
Long-context price tiers
Some models charge more per token once a prompt passes a size threshold, and the higher rate then applies to the whole request. GPT-6.1 Sol switches at 272,000 input tokens: a 269,000-token prompt with a 2,000-token answer costs $0.56, while a 281,000-token prompt costs $1.15. If your prompts sit near a threshold, trimming them below it is the cheapest optimisation available.
Common LLM cost estimation mistakes
- Counting only the user’s message. System prompts, tool schemas and retrieved context are often most of the input.
- Pricing one turn of a conversation. History is re-sent each turn, so price a whole conversation or a whole agent task.
- Ignoring reasoning tokens. A short visible answer can carry thousands of billed thinking tokens.
- Using one model’s token count for another. Tokenizers differ; Anthropic says its newer tokenizer (Claude 4.7 and later) produces about 30% more tokens for the same text than the previous one. What is a token explains why counts differ.
- Assuming every prompt hits the cache. Caches expire and need an identical prefix. Measure your hit rate from the usage fields.
- Forgetting failures. Timeouts, invalid JSON and retries are billed like any other request.
- Mixing up per-thousand and per-million prices. Current price lists quote per million tokens; older articles often quote per thousand.
FAQ
Questions people ask
How much does it cost to run a chatbot on an LLM API?
It depends mostly on the model and on message volume. Our chatbot preset (1,500 input tokens with half cached, 400 output, 1,000 messages a day) costs $169.57 a month on Claude Sonnet 5.5 and $21.58 on DeepSeek V4.1 Flash at today’s prices, or $0.00558 per message on Sonnet. Enter your own numbers in the cost calculator.
Why are output tokens more expensive than input tokens?
The model reads your whole prompt in parallel but writes its reply one token at a time, and each new token needs a full pass through the model. That makes output more expensive to serve. On popular models, output costs 3× to 5× the input price, so long replies and reasoning usually dominate the bill.
How many tokens does a conversation use?
More than the messages suggest, because every turn resends the system prompt and the full history. A 10-turn chat with a 1,500-token system prompt, 100-token messages and 300-token replies sends 34,000 input tokens in total, about 2.1 times what multiplying the first turn by 10 would predict.
Is the batch API worth it?
Yes, for any work that can wait. OpenAI, Anthropic and Google charge 50% less for batch requests, which are processed asynchronously: OpenAI promises completion within 24 hours and Anthropic says most batches finish within an hour. It doesn’t suit chat, autocomplete or anything a user is waiting on.
How do I know how many tokens my requests really use?
Every API response includes a usage object with input, output and cached token counts, and reasoning tokens where the model thinks. Log those numbers for a day of real traffic and average them. Before launch, count your full prompt with the token counter or a provider’s token-counting endpoint (Anthropic and Google both offer one, and Anthropic’s is free to use).
How accurate is an LLM cost estimate?
Input can be estimated closely, because you control the prompt. Output varies with the task, the model and its reasoning settings, so it’s the main source of error. Estimate from a sample of real runs, add a margin for retries and long requests, and correct the estimate against the provider’s usage dashboard after the first week.
Try it
Tools from this guide
Keep reading