Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Tokens & Costs

LLM API cost calculator for 340+ models

Estimate what an LLM API will cost per request, per day and per month, with prompt caching, batch pricing and cheaper alternatives. Free, and it runs in your browser.

$2 input · $10 output per 1M tokens · 1,000,000-token context

Start from a typical workload

Prompt, history and context

The model's reply

Share of each prompt read from the prompt cache at $0.10 per 1M.

Per month$212.92
Per request$0.007
Per day$7.00
Per year$2,555
Input per request
$0.003
Output per request
$0.004
    Prices updated 2026-10-09

    10 cheapest models for this workload

    Only models whose context and output limits fit your request
    ModelPer monthAction
    Mistral NemoMistral$1.69
    Ling 3.0 Flash VLinclusionAI$1.71
    Ling 3.0 FlashinclusionAI$1.72
    gpt-oss-20bOpenAI$1.92
    Granite 4.0 MicroIBM$2.14
    Nex-N2.5-MiniNex AGI$2.36
    Llama 3 8B LunarisSao10K$2.43
    Qwen3.7 FlashQwen$2.95
    Schematron V2 TurboInference.net$3.19
    Llama 3.1 8B InstructMeta$3.25

    Cheapest isn’t always best: quality, speed and reliability differ a lot between models. Use this list to find candidates, then test them on your own prompts.

    Steps

    How to use the LLM API cost calculator

    1. Pick the model you plan to use. Search by model name or provider.
    2. Enter the typical number of input and output tokens per request, or start from a preset such as a support chatbot or coding agent.
    3. Set how many requests you expect per day.
    4. If your prompts share a long, fixed start (like a system prompt), set the cached share. If results can wait up to a day, turn on the batch API.
    5. Read the monthly cost, compare it with the cheapest alternatives, and copy the link to share your estimate.

    Method

    How it works

    Every major LLM API bills by the token: small chunks of text that average about four characters in English. Input tokens (everything you send: system prompt, conversation history, retrieved documents) and output tokens (the model’s reply) have separate prices, quoted in US dollars per million tokens. This calculator applies those published prices to your workload.

    The formula

    For one request, the cost is:

    (uncached input × input price + cached input × cached price + output × output price) ÷ 1,000,000

    The daily cost multiplies that by your requests per day. The monthly figure uses the average month length of 365 ÷ 12 = 30.4 days, and the yearly figure uses 365 days.

    Output tokens are the expensive part

    Output is priced at 3× to 5× the input price for the popular models below, because each output token is generated one after another. Two things inflate output more than people expect. First, reasoning models bill their hidden “thinking” as output tokens, so a short visible answer can still cost thousands of tokens. Second, a verbose system prompt (“explain your reasoning step by step”) makes every reply longer. Asking for concise answers is often the cheapest optimisation available.

    Prompt caching

    If many requests begin with the same text, providers can keep that prefix in a cache and charge 2% to 27% of the normal input price to read it again. The calculator assumes the cache is already warm. Some providers also charge a one-time premium to write to the cache, which matters if your prefix changes often. Caching only applies to an identical prefix, so put the fixed parts of your prompt first and the parts that change last.

    Batch API and long-context pricing

    Batch APIs run requests in the background, usually within 24 hours, at roughly half price. When you turn on batch mode, the calculator uses each provider’s published batch price, and it tells you if a model has none. Some models also switch to a higher rate once a prompt passes a size threshold, such as 200,000 tokens. When that happens the higher rate covers the whole request, and the calculator shows a note.

    Choosing a cheaper model

    The alternatives table ranks every model whose context window and output limit can handle your request, from cheapest up. It doesn’t judge quality: the cheapest model may not be good enough for your task. Use it to shortlist candidates, then test them on real prompts before switching.

    Prices come from the OpenRouter models API and the open-source LiteLLM price list and are refreshed daily (last updated 2026-10-09). For open-weight models such as Llama, the price is a typical hosted price; running them on your own hardware costs something else entirely. The methodology page explains how the data is collected and checked.

    Examples

    Worked examples

    Support chatbot: 1,500 input and 400 output tokens, 1,000 conversations a day

    ModelPer requestPer monthPer month, half the prompt cached
    GPT-6.1 Sol OpenAI$0.007$212.92$169.57
    GPT-6 Luna OpenAI$0.00035$10.65$8.59
    Claude Opus 5.5 Anthropic$0.014$425.83$339.15
    Claude Sonnet 5.5 Anthropic$0.007$212.92$169.57
    Claude Haiku 5.5 Anthropic$0.00035$10.65$8.59
    Gemini 3.8 Flash Google$0.00262$79.84$64.45
    DeepSeek V4.1 Flash DeepSeek$0.00093$28.29$21.58
    Llama 4 Maverick Meta$0.000542$16.49$13.36

    Prices used, in USD per 1M tokens

    ModelInputCached inputOutputBatch in / outContext
    GPT-6.1 Sol$2$0.10$10$1 / $51,050,000
    GPT-6 Luna$0.10$0.01$0.50$0.05 / $0.251,050,000
    Claude Opus 5.5$4$0.20$20$2 / $101,000,000
    Claude Sonnet 5.5$2$0.10$10$1 / $51,000,000
    Claude Haiku 5.5$0.10$0.01$0.50$0.05 / $0.251,000,000
    Gemini 3.8 Flash$0.75$0.075$3.75$0.375 / $1.8751,048,576
    DeepSeek V4.1 Flash$0.30$0.006$1.20$0.112 / $0.3361,048,576
    Llama 4 Maverick$0.1875$0.05$0.6525–128,000

    Last updated 2026-10-09 from OpenRouter models API and LiteLLM model prices and context windows. “–” means no published price.

    FAQ

    Frequently asked questions

    How do I know how many tokens my requests use?

    For English text, one token is roughly four characters or three-quarters of a word, so 1,000 words is about 1,300 tokens. The most reliable way is to log the usage numbers the API returns with every response, average them over a day of real traffic, and enter those averages here.

    Why do output tokens cost more than input tokens?

    Input tokens are processed in parallel, but output tokens are generated one at a time, which uses far more compute per token. Across the popular models on this page, output is priced at 3× to 5× the input price, so long answers usually dominate the bill.

    What is cached input, and should I use it?

    When many requests start with the same text (a long system prompt, tool definitions, or a document you ask several questions about), providers can cache that prefix and charge much less to read it again: 2% to 27% of the normal input price for the models on this page. If your prompts share a stable prefix, set the cached share to roughly how much of each prompt that prefix makes up.

    When is the batch API worth it?

    Batch APIs process requests asynchronously, usually within 24 hours, for about half the normal price. They suit work nobody is waiting for: nightly summaries, classifying a backlog, generating embeddings or evaluations. They don't suit chat or anything interactive.

    Are these the prices I'll actually pay?

    They are list prices in US dollars per million tokens. Your bill can differ because of taxes, free tiers, volume discounts, regional pricing, or using a reseller instead of the provider directly. For open-weight models, the price shown is a typical hosted price, and different hosts charge different amounts. Always confirm critical numbers on the provider's pricing page.

    How often are the prices updated?

    Every day. A scheduled job pulls the OpenRouter models API and the open-source LiteLLM price list, checks the data, and publishes it if anything changed. The prices on this page were last updated on 2026-10-09. See the methodology for details.

    Why does the cost jump for very long prompts?

    Some models charge a higher rate once a single prompt passes a size threshold (often around 200,000 tokens), and the higher rate then applies to the whole request, not just the extra tokens. The calculator applies these long-context prices automatically and tells you when they kick in.