AI glossary · Models and context
What is a context window?
Also called: context length, context size, token limit
Definition
A context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.
Explained
How it works
Models read and write tokens, not words, and each request has a fixed budget of them. Everything you send counts: the system prompt, tool definitions, earlier turns, pasted documents and tool results. The reply counts too, including any hidden reasoning tokens, so the room you reserve with max output tokens has to fit as well.
The model keeps nothing between requests. A chat app re-sends the conversation every turn, so the window fills as you talk. If the input alone is too big, the API refuses it: OpenAI’s Responses API returns a 400 error by default, and Anthropic’s returns “prompt is too long”. Apps and agents stay inside the limit by dropping or summarising old turns, and the model then stops seeing what was removed.
Example
How many English words fit in a million-token window
Window sizes come from our daily data. To turn them into words we use measured ratios: 1,747 words of English from the Universal Declaration of Human Rights came to 2,017 tokens with OpenAI’s o200k_base tokenizer and 2,072 with Google’s token counter. Anthropic publishes its own figure for current Claude models.
Three windows of about a million tokens hold quite different amounts of text, because each tokenizer splits English differently. These are ceilings for plain prose: code, tables, JSON and most other languages use more tokens per word, and the reply has to fit in the same space. Check your own text with the context window checker.
| Model | Context window | Ratio used | About this many words |
|---|---|---|---|
| GPT-6.1 Sol | 1,050,000 tokens | 1,747 words = 2,017 tokens (o200k_base) | 909,000 |
| Gemini 3.8 Flash | 1,048,576 tokens | 1,747 words = 2,072 tokens (Gemini count) | 884,000 |
| Claude Sonnet 5.5 | 1,000,000 tokens | About 555,000 words per 1M tokens (Anthropic) | 555,000 |
Windows from our daily data, 2026-10-11. Ratios measured 2026-10-11 (the tokens to words converter has the full set) and from Anthropic’s models overview. Words rounded to the nearest thousand.
Cost and quality
Why it matters
The window caps what one request can see, but you pay for every token in it on every call, and long prompts take longer to start answering. Treat a big window as room, not a target: retrieve only the relevant passages with RAG, or cache a repeated prefix, rather than resending everything.
More context also isn’t free for quality. Anthropic’s docs say accuracy and recall degrade as the token count grows, which it calls “context rot”, so what you put in the window matters as much as how much fits.
Don’t mix up
Common confusions
- Context window vs max output tokens
- The window is the whole budget for one request. Max output tokens is a separate, much smaller cap on the reply. A model with a million-token window still stops when the reply reaches its output cap.
- Context window vs what the model knows
- What a model learned in training is not in the window, and nothing in the window is remembered after the request ends unless your app sends it again. Anthropic describes the window as the model’s “working memory”.
Go deeper
Try it and read more
- Free toolContext Window CheckerSee whether your text fits each model’s context window.
- Free toolTokens to Words ConverterConvert between tokens, words and pages.
- Free toolAI Model ComparisonCompare prices, context windows and features across models.
- Guide · 10 min readContext windows explained: what counts, what happens at the limit, and how to check fitWhat an LLM context window is, what counts toward it, max output vs context, what happens when you exceed it, and when RAG beats long context.
Related
Related terms
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- Max output tokensMax output tokens is the cap on how many tokens a model may generate in one response, set per request up to the model’s own limit, with any reasoning tokens counted inside it.
- RAGRAG (retrieval-augmented generation) is a technique where an application searches your own documents for passages relevant to a question and adds them to the prompt, so the model answers from that text rather than from memory alone.
- Prompt cachingPrompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.
- KV cacheThe KV cache is the memory where a language model keeps the attention keys and values of every token it has already processed, so each new token is computed without reprocessing the whole sequence.