Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

Tokens & Costs

Batch API calculator: real-time vs batch cost

Batch APIs from OpenAI, Anthropic and Google cost half the real-time price if you can wait for the results. This free calculator, which runs in your browser, shows what that saves on your workload using each model’s listed batch price, with turnaround and limits for each provider.

Real-time $2 / $10 per 1M input / output. Batch $1 / $5 (50% off).

Start from an example
Workload

Average daily volume

Instructions plus the item

The model’s reply

Requests nobody is waiting on. The rest stay real-time.

Share of each prompt read from the prompt cache ($0.10 per 1M in real time).

Saved per month$1,59750% less than running it all in real time
All real-time$3,194 per month
With 100% through batch$1,597 per month
Cost by period, real-time and with batch
PeriodReal-timeWith batchSaved
Per day$105.00$52.50$52.50
Per month$3,194$1,597$1,597
Per year$38,325$19,163$19,163
Per request$0.0021$0.00105$0.00105
  • 50,000 requests through batch, 0 in real time each day.
  • Anthropic Message Batches API: Most batches finish in less than 1 hour. Results are available once every request is done, or after 24 hours. Limits and expiry
Prices updated 2026-10-11

10 cheapest batch options for this workload

Models with a listed batch price whose limits fit your request
ModelPer monthAction
GPT-5 NanoOpenAI$45.63
gpt-oss-120bOpenAI · open weights, hosted price$46.36
Gemini 2.5 Flash LiteGoogle$76.04
GPT-4.1 NanoOpenAI$76.04
Claude Haiku 5.5Anthropic$79.84
GPT-6 LunaOpenAI$79.84
GPT-6 Luna ProOpenAI$79.84
GLM 5.3 FlashZ.ai · open weights, hosted price$88.21
Ministral 3 8B 2512Mistral · open weights, hosted price$96.95
GPT-4o-miniOpenAI$114.06

Costs use your split between batch and real time. Cheapest isn’t always best: quality differs a lot between models, so test candidates on your own prompts. Models without a published batch price are left out rather than given an assumed discount.

When batch is a good fit

Good fit: nobody is waiting

  • Evaluation runs and benchmark sweeps over a fixed test set
  • Classifying or tagging a backlog: tickets, reviews, transactions
  • Summarising or extracting data from a document archive
  • Nightly reports, enrichment and data clean-up jobs
  • Generating embeddings or synthetic data in bulk

Poor fit: someone needs the answer now

  • Chat, autocomplete or anything a person waits on
  • Agents that need each answer before the next step
  • Work with a deadline shorter than the provider’s window (24 hours at OpenAI)
  • Small one-off requests, where the setup outweighs the saving

Steps

How to use the batch API calculator

  1. Pick the model you plan to use. The line under it shows its real-time and batch prices, or says that it has no batch price.
  2. Choose a daily workload or a one-off job, then enter the number of requests and the input and output tokens per request. An example fills typical sizes.
  3. Set the share of requests that can wait for batch results. Anything a person waits on stays real-time.
  4. If every request starts with the same instructions, set the cached share to see caching and batch together.
  5. Read the saving, compare the cheapest batch options, and copy the link to share the estimate.

Method

How it works

A batch API takes a file of requests, runs them in the background when the provider has spare capacity, and returns a file of results. In exchange for waiting, you pay less: OpenAI, Anthropic and Google each charge 50% of their standard price for batch requests, on input and output alike. Batch traffic also has its own rate limits, so a big job doesn’t use up the quota your live app needs.

The formula

For each request, the calculator works out two prices from the model’s listed rates:

real-time = (input × input price + output × output price) ÷ 1,000,000
batch = (input × batch input price + output × batch output price) ÷ 1,000,000

Your workload is then split: the share that can wait is priced at the batch rate, the rest at the real-time rate. A daily cost becomes a monthly one over 365 ÷ 12 = 30.4 days. Long-context prices, where a model charges more once a prompt passes a threshold, are scaled by the same batch ratio.

We never assume a discount. Only 80 of the 321 text models in our data list a batch price; for the rest, the calculator says “no batch price published” and leaves them out of the ranking. A listed batch price above the real-time price is treated as a data error and ignored.

How the three providers differ

OpenAI’s Batch API has a single 24-hour completion window. Requests unfinished at the end are cancelled, and you pay only for completed ones. A batch holds up to 50,000 requests. Anthropic’s Message Batches API is usually the fastest: most batches finish in less than 1 hour. Results are available once every request is done, or after 24 hours. A batch holds up to 100,000 requests, and anything not processed within 24 hours expires unbilled. Google’s Gemini Batch API targets 24 hours and expires jobs after 48, and it is not available on the free tier. The cards under Worked examples have the full terms.

Batch plus prompt caching

Caching cuts the price of a repeated prompt start, such as shared instructions, and it works inside batches too. The calculator applies the batch ratio to the cached-input price (except on Gemini, where Google’s batch guide says cache hits pay the standard caching rate) and assumes the cache is already warm, so it leaves out cache-write charges. That can flatter the result: Anthropic says cache hits in batches are best-effort, typically 30% to 98%, because requests run concurrently in any order. Gemini batch requests need an explicit cache, which charges hourly storage. Some older OpenAI models list no batch cached-input price at all. To model writes, lifetimes and hit rates properly, use the prompt caching calculator.

Batch, flex and priority

OpenAI and Google also sell a middle option. OpenAI’s Flex processing (in beta, for a limited set of models) bills ordinary synchronous requests at Batch API rates, but responses are slower and can be refused when capacity is short. Google’s Flex inference is also 50% off, with a 1 to 15 minute latency target. At the other end, OpenAI’s Fast mode and Google’s Priority tier cost more than standard for lower latency.

Prices come from the OpenRouter models API and the LiteLLM price list, refreshed daily (last updated 2026-10-11). Provider terms were read from their documentation on 2026-10-11. To compare real-time prices across every model, use the model comparison table; for a full cost estimate with caching and long-context tiers, use the LLM cost calculator. The guides to estimating LLM API costs and the cheapest LLM APIs go further.

Examples

Worked examples

Batch APIs compared: OpenAI, Anthropic and Google

OpenAI

Batch API

Price
50% lower costs than the synchronous API.
Turnaround
A 24-hour completion window, which is currently the only option. OpenAI calls it “a clear 24-hour turnaround time”.
If it runs late
Requests still unfinished after the window are cancelled and written to the error file. You’re charged for completed requests, whose results stay in the output file.
Limits
Up to 50,000 requests and a 200 MB input file per batch; up to 2,000 batches created per hour; a per-model cap on queued prompt tokens. Works with the Responses, Chat Completions, Embeddings, Completions, Moderations and image endpoints.
Results kept
Output files are deleted 30 days after the batch completes.
Rate limits
Separate from the standard per-model rate limits, so batch work doesn’t use up your real-time quota.
With caching
OpenAI’s pricing page lists a batch cached-input price for most current models, at half the standard cached price. Some older models list none, so for them caching may not reduce a batch bill.
Other tiers
Flex processing (beta, limited models) bills ordinary synchronous requests at Batch API rates, with slower responses and occasional “429 Resource Unavailable” errors that aren’t charged. Fast mode (formerly Priority processing) costs more than standard.

Anthropic

Message Batches API

Price
All usage is charged at 50% of standard API prices.
Turnaround
Most batches finish in less than 1 hour. Results are available once every request is done, or after 24 hours.
If it runs late
A batch expires if processing doesn’t finish within 24 hours. Expired, cancelled and errored requests aren’t billed.
Limits
Up to 100,000 requests or 256 MB per batch, whichever comes first. Every active model is supported; streaming isn’t. Busy periods can slow processing and make more requests expire.
Results kept
Results can be downloaded for 29 days after the batch is created.
Rate limits
Batch rate limits cover API calls and the number of requests waiting to be processed. Batches can go slightly over a workspace’s spend limit.
With caching
Prompt caching works in batches and the two discounts stack, but cache hits are best-effort because requests run concurrently. Anthropic says users typically see hit rates of 30% to 98% and suggests the 1-hour cache for shared context.

Google

Gemini Batch API

Price
Priced at 50% of the standard interactive API cost.
Turnaround
A target turnaround of 24 hours, though Google says most jobs finish much sooner.
If it runs late
A job pending or running for more than 48 hours expires, with no results to retrieve.
Limits
Inline requests up to 20 MB in total, or a JSON Lines input file up to 2 GB. Works with generateContent requests and embeddings. Not available on the free tier.
Results kept
Results are kept for 6 weeks, then permanently deleted.
Rate limits
Separate from non-batch calls: up to 100 concurrent batch requests, 20 GB of file storage and a per-model cap on queued tokens.
With caching
Context caching works in batches: each request names an explicit cache. The batch guide says cache hits pay “the standard context caching rates”. The pricing page agrees for Gemini 2.5, 3 Flash Preview and 3.1 Pro, but lists a lower batch rate for some newer models such as Gemini 3.8 Flash. This calculator uses the standard rate, so it can overstate a Gemini batch bill slightly. Explicit caches also charge hourly storage, which this calculator doesn’t include.
Other tiers
Flex inference is also 50% off, as ordinary synchronous calls with a 1 to 15 minute latency target and best-effort availability. Priority costs 75% to 100% more than standard for low latency.

Read from each provider’s documentation on 2026-10-11 (links under Sources). Terms change, so check them before you plan a deadline around a batch.

Nightly ticket classification

50,000 support tickets a night: instructions plus one ticket in, a label and a short reason out.

ModelPer month, real-timePer month, batchSaved
GPT-6.1 Sol OpenAI$3,194$1,597$1,597
GPT-6 Luna OpenAI$159.69$79.84$79.84
Claude Opus 5.5 Anthropic$6,388$3,194$3,194
Claude Sonnet 5.5 Anthropic$3,194$1,597$1,597
Claude Haiku 5.5 Anthropic$159.69$79.84$79.84
Gemini 3.8 Flash Google$1,198$598.83$598.83
DeepSeek V4.1 Flash DeepSeek$456.25$161.82$294.43

The cheapest batch option for this workload today is GPT-5 Nano (OpenAI) at $45.63 a month. Classification needs a short answer, so a small model is often enough: test it on a few hundred labelled tickets first.

Summarise a document archive

A one-off job: 20,000 documents of about 6,000 tokens each, a 500-token summary for each.

ModelPer job, real-timePer job, batchSaved
GPT-6.1 Sol OpenAI$340.00$170.00$170.00
GPT-6 Luna OpenAI$17.00$8.50$8.50
Claude Opus 5.5 Anthropic$680.00$340.00$340.00
Claude Sonnet 5.5 Anthropic$340.00$170.00$170.00
Claude Haiku 5.5 Anthropic$17.00$8.50$8.50
Gemini 3.8 Flash Google$127.50$63.75$63.75
DeepSeek V4.1 Flash DeepSeek$48.00$16.80$31.20

Evaluation run

5,000 test cases with a shared 1,000-token rubric (cached) and a 1,500-token case, 800 tokens out.

ModelPer job, real-timePer job, batchSaved
GPT-6.1 Sol OpenAI$55.50$27.75$27.75
GPT-6 Luna OpenAI$2.80$1.40$1.40
Claude Opus 5.5 Anthropic$111.00$55.50$55.50
Claude Sonnet 5.5 Anthropic$55.50$27.75$27.75
Claude Haiku 5.5 Anthropic$2.80$1.40$1.40
Gemini 3.8 Flash Google$21.00$10.69$10.31
DeepSeek V4.1 Flash DeepSeek$7.08$2.20$4.88

Evaluation runs suit batch well: the test set is fixed, nobody waits on a single answer, and a shared rubric at the start of each prompt can be cached. On Anthropic models, treat the cached share as a best case, since batch cache hits are best-effort.

Computed from our price data on 2026-10-11. Every request is sent through batch in these examples; models without a published batch price are left out.

FAQ

Frequently asked questions

How much does the batch API save?

OpenAI, Anthropic and Google all charge 50% of their standard prices for batch requests, on both input and output tokens. In our data, 80 of 321 text models list a batch price. Listings from other providers and hosts show discounts from 20% to 72%. Your overall saving depends on how much of the workload can wait.

How long does a batch API job take?

OpenAI gives each batch a 24-hour completion window. Anthropic says most batches finish in less than an hour, with a 24-hour limit. Google targets 24 hours and says most jobs are much quicker, and expires jobs after 48 hours. Plan for the full window: a batch can take longer when the provider is busy.

Does prompt caching work with the batch API?

Yes, at all three providers, and the discounts stack. Anthropic warns that cache hits in batches are best-effort, typically 30% to 98%, and suggests the 1-hour cache for shared context. Gemini batch requests use an explicit cache, which charges hourly storage. OpenAI lists batch cached-input prices for most current models.

What happens if a batch doesn’t finish in time?

At OpenAI, unfinished requests are cancelled when the 24-hour window ends; you pay for completed ones and get their results. At Anthropic, requests not processed within 24 hours expire and aren’t billed. At Google, a job still pending or running after 48 hours expires with no results. Keep each request’s ID so you can resubmit what’s missing.

What are the batch API limits?

OpenAI allows up to 50,000 requests and a 200 MB file per batch, and up to 2,000 batches an hour. Anthropic allows 100,000 requests or 256 MB per batch. Google accepts inline requests up to 20 MB or an input file up to 2 GB, with up to 100 batches running at once. All three cap queued work, separately from real-time limits.

What is the difference between batch and flex processing?

Batch is asynchronous: you upload many requests and collect the results later. Flex is an ordinary synchronous request at a lower price that may be slower or turned away when capacity is short. OpenAI bills Flex at Batch API rates; Google’s Flex is 50% off with a 1 to 15 minute latency target. Both are best for work that isn’t urgent.