AI glossary · Tokens and cost
What is a rate limit in an AI API?
Also called: RPM and TPM limits, 429 Too Many Requests
Definition
A rate limit is a cap on how many requests or tokens an account may send to an API per minute or per day, and going over it makes the API reject requests with HTTP 429 until the allowance refills.
Explained
How it works
LLM APIs measure several limits at once and stop you at whichever you hit first. OpenAI counts requests and tokens per minute and per day (RPM, RPD, TPM, TPD). Anthropic splits tokens by direction: requests per minute (RPM), input tokens per minute (ITPM) and output tokens per minute (OTPM), per model. Both raise limits automatically through usage tiers as your account builds history and spend.
Limits are not reset once a minute. Anthropic uses a token bucket that refills continuously, and it may enforce 60 RPM as one request per second, so a short burst can fail. Responses carry headers showing what is left, such as x-ratelimit-remaining-tokens on OpenAI and anthropic-ratelimit-input-tokens-remaining on Anthropic. Go over and you get a 429, usually with a retry-after header giving the seconds to wait.
What counts differs. On most Claude models, tokens read from the prompt cache don’t count toward ITPM. OpenAI counts each request using the higher of max_tokens and an estimate of its size, so an oversized max_tokens uses up limit you never spend.
Example
How many requests a minute fit in Claude Sonnet 5.5’s Start tier
Anthropic’s standard Start tier allows Claude Sonnet 5.5 1,000 requests, 2,000,000 input tokens and 400,000 output tokens per minute (its rate limits page, 2026-10-11; new accounts may start lower). An agent that sends 30,000 input tokens and gets 800 back per request hits the input-token limit first, at 66 requests a minute, far below the request limit.
If 25,000 of those input tokens are read from the cache, only 5,000 count, and the same tier allows 400 requests a minute: 6 times the throughput with no change of tier.
| Limit | Allowance | No caching | 25,000 tokens cached |
|---|---|---|---|
| Requests per minute | 1,000 | 1,000 | 1,000 |
| Input tokens per minute | 2,000,000 | 66 | 400 |
| Output tokens per minute | 400,000 | 500 | 500 |
| Requests per minute you get | 66 | 400 |
Limits from Anthropic’s rate limits page on 2026-10-11. Tiers and limits change; check your own in the Claude Console.
Cost and quality
Why it matters
Rate limits, not price, are often what stops an agent or batch job from scaling. Parallel subagents and retries burn through tokens per minute quickly, and a wave of 429s can look like an outage.
Honour retry-after; without it, back off exponentially with random jitter. Don’t retry 429s that need action: one caused by Anthropic’s monthly spend cap has no retry-after and keeps failing. Move work that can wait to a batch API, which has its own limits. If you hit this in Claude Code, see Request rejected (429).
Don’t mix up
Common confusions
- Rate limit vs spend limit
- A rate limit caps speed and refills within a minute. A spend cap limits money per month. On Anthropic, hitting your tier’s monthly spend cap also returns a 429, so read the error:
enforced_spend_limit_reachedwon’t clear by retrying. A lower spend limit you set yourself returns a 400 instead. - 429 vs 529
- A 429 is about your account’s usage. Anthropic’s 529
overloaded_errormeans the API is busy for everyone; it isn’t fixed by a higher tier, only by waiting and retrying.
Go deeper
Try it and read more
Related
Related terms
- API keyAn API key is a secret string that identifies your account to a service such as the OpenAI, Claude or Gemini API, so every request made with it is authorised, rate-limited and billed to you.
- Batch APIA batch API is an asynchronous way to send many model requests as one job, which the provider works through when it has capacity, typically within 24 hours, at half the normal per-token price on OpenAI, Anthropic and Google.
- Prompt cachingPrompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.