AI glossary · Tokens and cost
What is a batch API?
Also called: Message Batches API, batch processing, batch mode
Definition
A batch API is an asynchronous way to send many model requests as one job, which the provider works through when it has capacity, typically within 24 hours, at half the normal per-token price on OpenAI, Anthropic and Google.
Explained
How it works
Instead of calling the API once per request and waiting, you submit a list of requests (a JSONL file with one request per line on OpenAI and Google, a JSON list in the request body on Anthropic), with your own custom_id on each. You poll the job and download the results when it ends. Results can come back in any order, which is why each one carries your ID.
OpenAI, Anthropic and Google all charge 50% of the standard price for batched work. OpenAI completes batches within 24 hours and Google targets 24 hours; Anthropic says most batches finish in under an hour and expire after 24. Size limits differ: up to 50,000 requests or 200 MB per batch on OpenAI, 100,000 requests or 256 MB on Anthropic.
Batches run on their own rate limits, separate from the real-time API, so a big overnight job doesn’t eat your live traffic’s headroom. Anthropic notes that the batch and prompt caching discounts can stack, though cache hits in a batch are best effort.
Example
10,000 support tickets labelled overnight
A team tags 10,000 support tickets each night: 1,500 input tokens per ticket (instructions plus the ticket) and a 200-token JSON label back. Nobody reads the labels before morning, so the job can wait.
On Gemini 3.8 Flash, the real-time price is $0.75 input and $3.75 output per million tokens; the batch price is $0.375 and $1.875. The nightly job costs $18.75 in real time and $9.38 as a batch, 50% less, or $281.25 saved over 30 nights. The batch API calculator runs the same sum for your workload.
| Real-time API | Batch API | |
|---|---|---|
| Input tokens | 15,000,000 | 15,000,000 |
| Output tokens | 2,000,000 | 2,000,000 |
| Results arrive | Seconds | Within 24 hours |
| Cost | $18.75 | $9.38 |
Prices from our daily data, 2026-10-11.
Cost and quality
Why it matters
For work that doesn’t need an answer in seconds, such as evaluations, classification, data extraction, embeddings or nightly reports, a batch API is the simplest way to halve the bill without changing the model or the prompt.
The price is latency and some extra code: you can’t stream, and you must handle failed and expired requests. Anything a person is waiting for still belongs on the real-time API.
Don’t mix up
Common confusions
- Batch API vs parallel requests
- Firing many requests at once from your own code still uses the real-time API at full price and counts against your normal rate limits. A batch API is a separate endpoint with its own limits and the discount.
- Batch API vs several questions in one prompt
- Packing ten questions into one prompt is a single request: it saves repeating the instructions but risks mixed-up answers. A batch keeps every request separate and gives each its own result.
Go deeper
Try it and read more
- Free toolBatch API Savings CalculatorCompare batch and real-time API costs.
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Guide · 9 min readThe cheapest LLM APIs right now, ranked from daily price dataThe cheapest LLM APIs ranked by blended price from daily data, the cheapest for tools, vision and long context, free tiers, and how to cut cost per task.
- Guide · 11 min readHow to estimate LLM API costs: the formula, worked examples and the trapsEstimate LLM API costs per request and per month: the token formula, worked examples for chatbots, RAG and agents, plus caching, batch and reasoning tokens.
Related
Related terms
- Prompt cachingPrompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.
- Rate limitA rate limit is a cap on how many requests or tokens an account may send to an API per minute or per day, and going over it makes the API reject requests with HTTP 429 until the allowance refills.
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- Reasoning tokensReasoning tokens are the tokens a reasoning model generates while it works through a problem before writing its answer, billed as output tokens even though the API hides them or returns only a summary.