Tokens & Costs
Batch API calculator: real-time vs batch cost
Batch APIs from OpenAI, Anthropic and Google cost half the real-time price if you can wait for the results. This free calculator, which runs in your browser, shows what that saves on your workload using each model’s listed batch price, with turnaround and limits for each provider.
Real-time $2 / $10 per 1M input / output. Batch $1 / $5 (50% off).
Average daily volume
Instructions plus the item
The model’s reply
Requests nobody is waiting on. The rest stay real-time.
Share of each prompt read from the prompt cache ($0.10 per 1M in real time).
| Period | Real-time | With batch | Saved |
|---|---|---|---|
| Per day | $105.00 | $52.50 | $52.50 |
| Per month | $3,194 | $1,597 | $1,597 |
| Per year | $38,325 | $19,163 | $19,163 |
| Per request | $0.0021 | $0.00105 | $0.00105 |
- 50,000 requests through batch, 0 in real time each day.
- Anthropic Message Batches API: Most batches finish in less than 1 hour. Results are available once every request is done, or after 24 hours. Limits and expiry
10 cheapest batch options for this workload
Models with a listed batch price whose limits fit your request| # | Model | Per month | Saved | Turnaround | Action |
|---|---|---|---|---|---|
| 1 | GPT-5 NanoOpenAI | $45.63 | 50% | Within 24 hours | |
| 2 | gpt-oss-120bOpenAI · open weights, hosted price | $46.36 | 20% | Depends on host | |
| 3 | Gemini 2.5 Flash LiteGoogle | $76.04 | 50% | Target 24 hours | |
| 4 | GPT-4.1 NanoOpenAI | $76.04 | 50% | Within 24 hours | |
| 5 | Claude Haiku 5.5Anthropic | $79.84 | 50% | Most within 1 hour | |
| 6 | GPT-6 LunaOpenAI | $79.84 | 50% | Within 24 hours | |
| 7 | GPT-6 Luna ProOpenAI | $79.84 | 50% | Within 24 hours | |
| 8 | GLM 5.3 FlashZ.ai · open weights, hosted price | $88.21 | 60% | Depends on host | |
| 9 | Ministral 3 8B 2512Mistral · open weights, hosted price | $96.95 | 50% | Depends on host | |
| 10 | GPT-4o-miniOpenAI | $114.06 | 50% | Within 24 hours |
Costs use your split between batch and real time. Cheapest isn’t always best: quality differs a lot between models, so test candidates on your own prompts. Models without a published batch price are left out rather than given an assumed discount.
When batch is a good fit
Good fit: nobody is waiting
- Evaluation runs and benchmark sweeps over a fixed test set
- Classifying or tagging a backlog: tickets, reviews, transactions
- Summarising or extracting data from a document archive
- Nightly reports, enrichment and data clean-up jobs
- Generating embeddings or synthetic data in bulk
Poor fit: someone needs the answer now
- Chat, autocomplete or anything a person waits on
- Agents that need each answer before the next step
- Work with a deadline shorter than the provider’s window (24 hours at OpenAI)
- Small one-off requests, where the setup outweighs the saving
Steps
How to use the batch API calculator
- Pick the model you plan to use. The line under it shows its real-time and batch prices, or says that it has no batch price.
- Choose a daily workload or a one-off job, then enter the number of requests and the input and output tokens per request. An example fills typical sizes.
- Set the share of requests that can wait for batch results. Anything a person waits on stays real-time.
- If every request starts with the same instructions, set the cached share to see caching and batch together.
- Read the saving, compare the cheapest batch options, and copy the link to share the estimate.
Method
How it works
A batch API takes a file of requests, runs them in the background when the provider has spare capacity, and returns a file of results. In exchange for waiting, you pay less: OpenAI, Anthropic and Google each charge 50% of their standard price for batch requests, on input and output alike. Batch traffic also has its own rate limits, so a big job doesn’t use up the quota your live app needs.
The formula
For each request, the calculator works out two prices from the model’s listed rates:
real-time = (input × input price + output × output price) ÷ 1,000,000batch = (input × batch input price + output × batch output price) ÷ 1,000,000
Your workload is then split: the share that can wait is priced at the batch rate, the rest at the real-time rate. A daily cost becomes a monthly one over 365 ÷ 12 = 30.4 days. Long-context prices, where a model charges more once a prompt passes a threshold, are scaled by the same batch ratio.
We never assume a discount. Only 80 of the 321 text models in our data list a batch price; for the rest, the calculator says “no batch price published” and leaves them out of the ranking. A listed batch price above the real-time price is treated as a data error and ignored.
How the three providers differ
OpenAI’s Batch API has a single 24-hour completion window. Requests unfinished at the end are cancelled, and you pay only for completed ones. A batch holds up to 50,000 requests. Anthropic’s Message Batches API is usually the fastest: most batches finish in less than 1 hour. Results are available once every request is done, or after 24 hours. A batch holds up to 100,000 requests, and anything not processed within 24 hours expires unbilled. Google’s Gemini Batch API targets 24 hours and expires jobs after 48, and it is not available on the free tier. The cards under Worked examples have the full terms.
Batch plus prompt caching
Caching cuts the price of a repeated prompt start, such as shared instructions, and it works inside batches too. The calculator applies the batch ratio to the cached-input price (except on Gemini, where Google’s batch guide says cache hits pay the standard caching rate) and assumes the cache is already warm, so it leaves out cache-write charges. That can flatter the result: Anthropic says cache hits in batches are best-effort, typically 30% to 98%, because requests run concurrently in any order. Gemini batch requests need an explicit cache, which charges hourly storage. Some older OpenAI models list no batch cached-input price at all. To model writes, lifetimes and hit rates properly, use the prompt caching calculator.
Batch, flex and priority
OpenAI and Google also sell a middle option. OpenAI’s Flex processing (in beta, for a limited set of models) bills ordinary synchronous requests at Batch API rates, but responses are slower and can be refused when capacity is short. Google’s Flex inference is also 50% off, with a 1 to 15 minute latency target. At the other end, OpenAI’s Fast mode and Google’s Priority tier cost more than standard for lower latency.
Prices come from the OpenRouter models API and the LiteLLM price list, refreshed daily (last updated 2026-10-11). Provider terms were read from their documentation on 2026-10-11. To compare real-time prices across every model, use the model comparison table; for a full cost estimate with caching and long-context tiers, use the LLM cost calculator. The guides to estimating LLM API costs and the cheapest LLM APIs go further.
Examples
Worked examples
Batch APIs compared: OpenAI, Anthropic and Google
OpenAI
Batch API
- Price
- 50% lower costs than the synchronous API.
- Turnaround
- A 24-hour completion window, which is currently the only option. OpenAI calls it “a clear 24-hour turnaround time”.
- If it runs late
- Requests still unfinished after the window are cancelled and written to the error file. You’re charged for completed requests, whose results stay in the output file.
- Limits
- Up to 50,000 requests and a 200 MB input file per batch; up to 2,000 batches created per hour; a per-model cap on queued prompt tokens. Works with the Responses, Chat Completions, Embeddings, Completions, Moderations and image endpoints.
- Results kept
- Output files are deleted 30 days after the batch completes.
- Rate limits
- Separate from the standard per-model rate limits, so batch work doesn’t use up your real-time quota.
- With caching
- OpenAI’s pricing page lists a batch cached-input price for most current models, at half the standard cached price. Some older models list none, so for them caching may not reduce a batch bill.
- Other tiers
- Flex processing (beta, limited models) bills ordinary synchronous requests at Batch API rates, with slower responses and occasional “429 Resource Unavailable” errors that aren’t charged. Fast mode (formerly Priority processing) costs more than standard.
Anthropic
Message Batches API
- Price
- All usage is charged at 50% of standard API prices.
- Turnaround
- Most batches finish in less than 1 hour. Results are available once every request is done, or after 24 hours.
- If it runs late
- A batch expires if processing doesn’t finish within 24 hours. Expired, cancelled and errored requests aren’t billed.
- Limits
- Up to 100,000 requests or 256 MB per batch, whichever comes first. Every active model is supported; streaming isn’t. Busy periods can slow processing and make more requests expire.
- Results kept
- Results can be downloaded for 29 days after the batch is created.
- Rate limits
- Batch rate limits cover API calls and the number of requests waiting to be processed. Batches can go slightly over a workspace’s spend limit.
- With caching
- Prompt caching works in batches and the two discounts stack, but cache hits are best-effort because requests run concurrently. Anthropic says users typically see hit rates of 30% to 98% and suggests the 1-hour cache for shared context.
Gemini Batch API
- Price
- Priced at 50% of the standard interactive API cost.
- Turnaround
- A target turnaround of 24 hours, though Google says most jobs finish much sooner.
- If it runs late
- A job pending or running for more than 48 hours expires, with no results to retrieve.
- Limits
- Inline requests up to 20 MB in total, or a JSON Lines input file up to 2 GB. Works with generateContent requests and embeddings. Not available on the free tier.
- Results kept
- Results are kept for 6 weeks, then permanently deleted.
- Rate limits
- Separate from non-batch calls: up to 100 concurrent batch requests, 20 GB of file storage and a per-model cap on queued tokens.
- With caching
- Context caching works in batches: each request names an explicit cache. The batch guide says cache hits pay “the standard context caching rates”. The pricing page agrees for Gemini 2.5, 3 Flash Preview and 3.1 Pro, but lists a lower batch rate for some newer models such as Gemini 3.8 Flash. This calculator uses the standard rate, so it can overstate a Gemini batch bill slightly. Explicit caches also charge hourly storage, which this calculator doesn’t include.
- Other tiers
- Flex inference is also 50% off, as ordinary synchronous calls with a 1 to 15 minute latency target and best-effort availability. Priority costs 75% to 100% more than standard for low latency.
Read from each provider’s documentation on 2026-10-11 (links under Sources). Terms change, so check them before you plan a deadline around a batch.
Nightly ticket classification
50,000 support tickets a night: instructions plus one ticket in, a label and a short reason out.
| Model | Per month, real-time | Per month, batch | Saved |
|---|---|---|---|
| GPT-6.1 Sol OpenAI | $3,194 | $1,597 | $1,597 |
| GPT-6 Luna OpenAI | $159.69 | $79.84 | $79.84 |
| Claude Opus 5.5 Anthropic | $6,388 | $3,194 | $3,194 |
| Claude Sonnet 5.5 Anthropic | $3,194 | $1,597 | $1,597 |
| Claude Haiku 5.5 Anthropic | $159.69 | $79.84 | $79.84 |
| Gemini 3.8 Flash Google | $1,198 | $598.83 | $598.83 |
| DeepSeek V4.1 Flash DeepSeek | $456.25 | $161.82 | $294.43 |
The cheapest batch option for this workload today is GPT-5 Nano (OpenAI) at $45.63 a month. Classification needs a short answer, so a small model is often enough: test it on a few hundred labelled tickets first.
Summarise a document archive
A one-off job: 20,000 documents of about 6,000 tokens each, a 500-token summary for each.
| Model | Per job, real-time | Per job, batch | Saved |
|---|---|---|---|
| GPT-6.1 Sol OpenAI | $340.00 | $170.00 | $170.00 |
| GPT-6 Luna OpenAI | $17.00 | $8.50 | $8.50 |
| Claude Opus 5.5 Anthropic | $680.00 | $340.00 | $340.00 |
| Claude Sonnet 5.5 Anthropic | $340.00 | $170.00 | $170.00 |
| Claude Haiku 5.5 Anthropic | $17.00 | $8.50 | $8.50 |
| Gemini 3.8 Flash Google | $127.50 | $63.75 | $63.75 |
| DeepSeek V4.1 Flash DeepSeek | $48.00 | $16.80 | $31.20 |
Evaluation run
5,000 test cases with a shared 1,000-token rubric (cached) and a 1,500-token case, 800 tokens out.
| Model | Per job, real-time | Per job, batch | Saved |
|---|---|---|---|
| GPT-6.1 Sol OpenAI | $55.50 | $27.75 | $27.75 |
| GPT-6 Luna OpenAI | $2.80 | $1.40 | $1.40 |
| Claude Opus 5.5 Anthropic | $111.00 | $55.50 | $55.50 |
| Claude Sonnet 5.5 Anthropic | $55.50 | $27.75 | $27.75 |
| Claude Haiku 5.5 Anthropic | $2.80 | $1.40 | $1.40 |
| Gemini 3.8 Flash Google | $21.00 | $10.69 | $10.31 |
| DeepSeek V4.1 Flash DeepSeek | $7.08 | $2.20 | $4.88 |
Evaluation runs suit batch well: the test set is fixed, nobody waits on a single answer, and a shared rubric at the start of each prompt can be cached. On Anthropic models, treat the cached share as a best case, since batch cache hits are best-effort.
Computed from our price data on 2026-10-11. Every request is sent through batch in these examples; models without a published batch price are left out.
FAQ
Frequently asked questions
How much does the batch API save?
OpenAI, Anthropic and Google all charge 50% of their standard prices for batch requests, on both input and output tokens. In our data, 80 of 321 text models list a batch price. Listings from other providers and hosts show discounts from 20% to 72%. Your overall saving depends on how much of the workload can wait.
How long does a batch API job take?
OpenAI gives each batch a 24-hour completion window. Anthropic says most batches finish in less than an hour, with a 24-hour limit. Google targets 24 hours and says most jobs are much quicker, and expires jobs after 48 hours. Plan for the full window: a batch can take longer when the provider is busy.
Does prompt caching work with the batch API?
Yes, at all three providers, and the discounts stack. Anthropic warns that cache hits in batches are best-effort, typically 30% to 98%, and suggests the 1-hour cache for shared context. Gemini batch requests use an explicit cache, which charges hourly storage. OpenAI lists batch cached-input prices for most current models.
What happens if a batch doesn’t finish in time?
At OpenAI, unfinished requests are cancelled when the 24-hour window ends; you pay for completed ones and get their results. At Anthropic, requests not processed within 24 hours expire and aren’t billed. At Google, a job still pending or running after 48 hours expires with no results. Keep each request’s ID so you can resubmit what’s missing.
What are the batch API limits?
OpenAI allows up to 50,000 requests and a 200 MB file per batch, and up to 2,000 batches an hour. Anthropic allows 100,000 requests or 256 MB per batch. Google accepts inline requests up to 20 MB or an input file up to 2 GB, with up to 100 batches running at once. All three cap queued work, separately from real-time limits.
What is the difference between batch and flex processing?
Batch is asynchronous: you upload many requests and collect the results later. Flex is an ordinary synchronous request at a lower price that may be slower or turned away when capacity is short. OpenAI bills Flex at Batch API rates; Google’s Flex is 50% off with a 1 to 15 minute latency target. Both are best for work that isn’t urgent.
Related
Related tools
- LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Prompt Caching CalculatorEstimate savings from prompt caching.
- AI Model ComparisonCompare prices, context windows and features across models.
- AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- Context Window CheckerSee whether your text fits each model’s context window.
- Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.