Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Tokens and cost

What is a rate limit in an AI API?

Also called: RPM and TPM limits, 429 Too Many Requests

Definition

A rate limit is a cap on how many requests or tokens an account may send to an API per minute or per day, and going over it makes the API reject requests with HTTP 429 until the allowance refills.

Explained

How it works

LLM APIs measure several limits at once and stop you at whichever you hit first. OpenAI counts requests and tokens per minute and per day (RPM, RPD, TPM, TPD). Anthropic splits tokens by direction: requests per minute (RPM), input tokens per minute (ITPM) and output tokens per minute (OTPM), per model. Both raise limits automatically through usage tiers as your account builds history and spend.

Limits are not reset once a minute. Anthropic uses a token bucket that refills continuously, and it may enforce 60 RPM as one request per second, so a short burst can fail. Responses carry headers showing what is left, such as x-ratelimit-remaining-tokens on OpenAI and anthropic-ratelimit-input-tokens-remaining on Anthropic. Go over and you get a 429, usually with a retry-after header giving the seconds to wait.

What counts differs. On most Claude models, tokens read from the prompt cache don’t count toward ITPM. OpenAI counts each request using the higher of max_tokens and an estimate of its size, so an oversized max_tokens uses up limit you never spend.

Example

How many requests a minute fit in Claude Sonnet 5.5’s Start tier

Anthropic’s standard Start tier allows Claude Sonnet 5.5 1,000 requests, 2,000,000 input tokens and 400,000 output tokens per minute (its rate limits page, 2026-10-11; new accounts may start lower). An agent that sends 30,000 input tokens and gets 800 back per request hits the input-token limit first, at 66 requests a minute, far below the request limit.

If 25,000 of those input tokens are read from the cache, only 5,000 count, and the same tier allows 400 requests a minute: 6 times the throughput with no change of tier.

Requests per minute each limit allows, Claude Sonnet 5.5, Start tier
LimitAllowanceNo caching25,000 tokens cached
Requests per minute1,0001,0001,000
Input tokens per minute2,000,00066400
Output tokens per minute400,000500500
Requests per minute you get66400

Limits from Anthropic’s rate limits page on 2026-10-11. Tiers and limits change; check your own in the Claude Console.

Cost and quality

Why it matters

Rate limits, not price, are often what stops an agent or batch job from scaling. Parallel subagents and retries burn through tokens per minute quickly, and a wave of 429s can look like an outage.

Honour retry-after; without it, back off exponentially with random jitter. Don’t retry 429s that need action: one caused by Anthropic’s monthly spend cap has no retry-after and keeps failing. Move work that can wait to a batch API, which has its own limits. If you hit this in Claude Code, see Request rejected (429).

Don’t mix up

Common confusions

Rate limit vs spend limit
A rate limit caps speed and refills within a minute. A spend cap limits money per month. On Anthropic, hitting your tier’s monthly spend cap also returns a 429, so read the error: enforced_spend_limit_reached won’t clear by retrying. A lower spend limit you set yourself returns a 400 instead.
429 vs 529
A 429 is about your account’s usage. Anthropic’s 529 overloaded_error means the API is busy for everyone; it isn’t fixed by a higher tier, only by waiting and retrying.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary

Written by Tahir Nazir. Checked .

How this was checked: Start-tier limits, cache-aware ITPM, the token bucket, header names, the spend-cap 429 and 529 checked against Anthropic’s rate limits and errors pages; RPM/TPM measures, header names, max_tokens counting and retry advice against OpenAI’s rate limits guide, all on 2026-10-11. Requests per minute computed from those limits.