Guide · Tokens & costs
Context windows explained: what counts, what happens at the limit, and how to check fit
A context window is the most tokens a model can handle in one request: the system prompt, tool definitions, the conversation so far, any documents, and the reply it writes, all sharing one budget. If the input alone is too big, the request fails; if the reply runs out of room, it gets cut off.
By Tahir NazirUpdated 10 min read
On this page
- What is a context window?
- What counts toward the context window?
- Context window vs max output tokens
- What happens when you exceed the context window?
- Is a bigger context window better?
- RAG or long context: which should you use?
- Will it fit? A long PDF, a codebase and a chat history
- How to check whether your text fits
- Questions people ask
What is a context window?
The context window is everything a model can look at while it writes one response, including the response itself. Anthropic calls it the model’s “working memory”: it is separate from what the model learned in training, and it’s measured in tokens, not words or characters.
Window sizes vary a lot. Of the 333 text models in our data today, 102 accept at least 1,000,000 tokens and 120 accept fewer than 200,000. Here are some current models, with the separate limit on how much each can write in one reply:
| Model | Context window | Max output per reply |
|---|---|---|
| Claude Sonnet 5.5 | 1,000,000 tokens | 128,000 tokens |
| Claude Haiku 5.5 | 1,000,000 tokens | 128,000 tokens |
| GPT-6.1 Sol | 1,050,000 tokens | 128,000 tokens |
| GPT-6 Luna | 1,050,000 tokens | 128,000 tokens |
| Gemini 3.8 Flash | 1,048,576 tokens | 65,536 tokens |
| Gemini 3.5 Flash Lite | 1,048,576 tokens | 65,536 tokens |
| Llama 4 Maverick | 128,000 tokens | 16,384 tokens |
From our daily data (OpenRouter models API and LiteLLM). The model comparison table lists every model.
How much text is that? Anthropic says 1M tokens is roughly 555,000 words on the tokenizer Claude has used since Opus 4.7, and about 750,000 words on earlier Claude models. Code, tables and non-English text use more tokens per word, so measure your own material.
What counts toward the context window?
Everything in the request, plus everything the model writes back. Anthropic’s documentation lists it plainly: the system prompt, every message (including tool results, images and documents), your tool definitions, and the output, including thinking. OpenAI describes the window the same way: input, output and, on reasoning models, reasoning tokens.
GPT-6.1 Sol · context window
1,050,000 tokens
- System prompt3,000 · 0.3%
- Tool definitions15,000 · 1%
- Conversation history60,000 · 6%
- Documents and files400,000 · 38%
- Room kept for reasoning and the reply25,000 · 2%
- Still free547,000 · 52%
Example request · window from our daily data · tokens · share of window
Two things catch people out. First, the room for the reply: OpenAI suggests reserving at least 25,000 tokens for reasoning and output when you start with its reasoning models. Second, the hidden parts. Tool definitions are re-sent with every request, and connecting many MCP servers can add thousands of tokens before you type anything. The example’s 478,000-token prompt is also past GPT-6.1 Sol’s 272,000-token long-context threshold, so it would be billed at the higher rate.
- Earlier turns are re-sent in full each time, so a conversation grows with every message.
- Thinking from earlier turns stays in context on current Claude models and counts as input; older models strip it automatically.
- Cached tokens still count. Prompt caching changes what you pay for a repeated prefix, not whether it takes up room.
Context window vs max output tokens
The context window is the total budget for one request. Max output is a separate, usually much smaller cap on how many tokens the model can generate in one reply, and it includes any thinking. A model can have a million-token window and still stop after its output limit. Where a provider publishes no separate cap, the reply still has to fit in what’s left of the window.
You set the reply limit yourself on each request (max_tokens on Claude, max_output_tokens on OpenAI), up to the model’s maximum. If the model reaches it, the reply ends early: OpenAI marks the response incomplete with the reason max_output_tokens, and Claude stops with stop_reason: "max_tokens". With a reasoning model, a low limit can be used up by thinking before any answer appears.
What happens when you exceed the context window?
It depends on which part overflows and on who manages the conversation.
- The input is too big. Anthropic’s API rejects it with a 400
invalid_request_error, “prompt is too long”, on every model. Nothing is generated. - The reply runs out of room. On Claude 4.5 models and newer, a request whose input plus
max_tokensexceeds the window is accepted, and generation stops withstop_reason: "model_context_window_exceeded"if it reaches the limit. OpenAI warns that tokens beyond the window “may be truncated”. - In chat apps and agents, the app usually handles it: it drops the oldest messages or summarises them. That keeps the conversation going, but the model quietly stops seeing whatever was dropped.
Summarising old turns is called compaction. Anthropic and OpenAI both offer it on the server: Claude replaces older turns with a summary it writes itself (in beta), and OpenAI’s Responses API can compact once a token threshold you set is crossed. In Claude Code, /compact does the same by hand; see our fix for “Prompt is too long”.
Is a bigger context window better?
Bigger lets you send more, but sending more has three costs, so use the space you need rather than all of it.
- Money. You pay for every input token on every request. On Claude Sonnet 5.5, one request with a 900,000-token prompt and a 1,000-token answer costs $1.81, against $0.05 with a 20,000-token prompt (prices as of 2026-10-09). Caching a repeated prefix brings the big one down to about $0.12, and some models charge a higher rate for long prompts.
- Speed. Google’s long-context guide notes that longer queries generally have higher latency, measured as time to the first token.
- Recall. Models don’t use every part of a long context equally well. Anthropic’s docs say accuracy and recall degrade as the token count grows, which it calls “context rot”, and Google says accuracy drops when you look for several pieces of information at once rather than one.
The best-known study is “Lost in the Middle” (Liu and others; posted in 2023, published in TACL in 2024). Testing multi-document question answering and key-value lookup, it found performance was often highest when the relevant information sat at the beginning or end of the input, and dropped significantly when it was in the middle, even for models built for long contexts. The models tested are now old, so treat it as a reason to test your own task, not as a current score.
RAG or long context: which should you use?
Use long context when the model needs to read the whole thing at once, and retrieval when it only needs the relevant parts. RAG (retrieval-augmented generation) means searching your documents first and sending only the best-matching passages with the question.
- Long context suits one-off analysis of a single document, questions that need the whole text (a summary, a contract review, “what changed between these two versions”), and corpora small enough to fit and cache.
- RAG suits collections bigger than any window, many questions against the same material, content that changes often, and apps that must answer quickly and cheaply or cite their sources.
- Both together is common: retrieve generously, then let a long-context model read twenty passages instead of three.
Chunking: how to split what doesn’t fit
Whether you retrieve or process a document in parts, you need chunks. Split on the document’s own structure (sections, headings, functions) rather than at a fixed character count, keep each chunk understandable on its own, and attach where it came from (title, section, page) so answers can cite it. A small overlap between neighbouring chunks stops a sentence that spans a boundary from being lost. For map-style jobs like summarising a long report, process each chunk, then combine the partial results in a final request.
Will it fit? A long PDF, a codebase and a chat history
Here are three common jobs, checked against three real windows with 8,192 tokens kept for the reply and 3,000 for a system prompt:
| Job | Tokens | Claude Sonnet 5.5 (1,000,000) | Mistral Small 4 (262,144) | Llama 4 Maverick (128,000) |
|---|---|---|---|---|
| A 200-page PDF (text only) | 300,000–600,000 | Fits (61% of the window) | Too large (split into 3 parts) | Too large (split into 6 parts) |
| A 50,000-line codebase | about 560,000 | Fits (57% of the window) | Too large (split into 3 parts) | Too large (split into 5 parts) |
| A 150-turn chat | 150,000 | Fits (16% of the window) | Fits (61% of the window) | Too large (split into 2 parts) |
Token counts are estimates: the PDF uses Anthropic’s per-page range, the codebase our measured lines-to-tokens ratio, the chat 1,000 tokens per turn. Parts are the number of pieces the text would need, each sent with the same system prompt and reply room. Verdicts use the same rules as our context window checker.
- A long PDF. Anthropic says each PDF page typically uses 1,500 to 3,000 tokens of text, and each page is also sent as an image, which costs more tokens on top. Claude accepts up to 600 pages per request (100 when the window is under 1M), so a big report can hit the page limit before the token limit.
- A codebase. We counted this site’s own code: 20,062 lines of TypeScript came to 224,859 tokens with OpenAI’s o200k_base tokenizer, about 11 tokens per line. Coding agents usually don’t send a whole repository; they search for and read the files a task needs, which is cheaper and keeps the context focused.
- A chat history. At about 1,000 tokens per turn (a question and its answer), Llama 4 Maverick’s window holds roughly 116 turns with the same reply room. Long before that, it’s worth summarising older turns, because each new message resends all of them.
How to check whether your text fits
Count the tokens, add what else goes in the request, and leave room for the reply. Our token counter counts text in your browser. For the closest count, use the provider’s own counter: Anthropic’s token counting endpoint (free to call) and Gemini’s countTokens method take the same request you would send, and OpenAI’s tokenizers are published for its older models. For a quick check across many models at once, use our checker:
If the request fits but is large, check the price too: the LLM cost calculator shows what a long prompt costs per day and month, and the model comparison table lets you filter models by minimum context window.
FAQ
Questions people ask
How many words fit in a 1-million-token context window?
Roughly 555,000 to 750,000 words of English, depending on the tokenizer. Anthropic gives about 555,000 words for current Claude models and about 750,000 for earlier ones, and OpenAI’s rule of thumb is three quarters of a word per token. Code and non-English text fit fewer words, so count your own text.
Does the model’s reply count toward the context window?
Yes. The reply, including any hidden thinking or reasoning tokens, is generated inside the same window as your prompt. That’s why you need to leave room for it: a prompt that fills the window leaves nothing for the answer.
What is the difference between context window and context length?
They mean the same thing: the maximum number of tokens a model can handle in one request, counting both input and output. You will also see “context size”. Max output is different: a separate, usually much smaller limit on how long a single reply can be.
Does a model remember earlier conversations?
Only what is in the current request. Chat apps and stored-conversation APIs resend the earlier turns for you, and those turns count toward the window. Where an app has a “memory” feature, it saves notes outside the conversation and adds the relevant ones back into the prompt, where they count too.
Do cached tokens count toward the context window?
Yes. Prompt caching makes a repeated prefix cheaper and faster to process, but those tokens are still part of the request and still take up room in the window. Anthropic’s documentation states this directly: caching changes what you pay, not whether the tokens count.
Is a model with a bigger context window slower?
Not because of the window size itself, but because of how much you send. Longer prompts take longer to process before the first token of the reply appears, as Google’s long-context guide notes. Sending only what the task needs keeps responses faster and cheaper.
Try it
Tools from this guide
- Context Window CheckerSee whether your text fits each model's context window.
- AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- AI Model ComparisonCompare prices, context windows and features across models.
- LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
Keep reading