Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Models and context

What is a context window?

Also called: context length, context size, token limit

Definition

A context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.

Explained

How it works

Models read and write tokens, not words, and each request has a fixed budget of them. Everything you send counts: the system prompt, tool definitions, earlier turns, pasted documents and tool results. The reply counts too, including any hidden reasoning tokens, so the room you reserve with max output tokens has to fit as well.

The model keeps nothing between requests. A chat app re-sends the conversation every turn, so the window fills as you talk. If the input alone is too big, the API refuses it: OpenAI’s Responses API returns a 400 error by default, and Anthropic’s returns “prompt is too long”. Apps and agents stay inside the limit by dropping or summarising old turns, and the model then stops seeing what was removed.

Example

How many English words fit in a million-token window

Window sizes come from our daily data. To turn them into words we use measured ratios: 1,747 words of English from the Universal Declaration of Human Rights came to 2,017 tokens with OpenAI’s o200k_base tokenizer and 2,072 with Google’s token counter. Anthropic publishes its own figure for current Claude models.

Three windows of about a million tokens hold quite different amounts of text, because each tokenizer splits English differently. These are ceilings for plain prose: code, tables, JSON and most other languages use more tokens per word, and the reply has to fit in the same space. Check your own text with the context window checker.

English prose that fits in the whole window
ModelContext windowRatio usedAbout this many words
GPT-6.1 Sol1,050,000 tokens1,747 words = 2,017 tokens (o200k_base)909,000
Gemini 3.8 Flash1,048,576 tokens1,747 words = 2,072 tokens (Gemini count)884,000
Claude Sonnet 5.51,000,000 tokensAbout 555,000 words per 1M tokens (Anthropic)555,000

Windows from our daily data, 2026-10-11. Ratios measured 2026-10-11 (the tokens to words converter has the full set) and from Anthropic’s models overview. Words rounded to the nearest thousand.

Cost and quality

Why it matters

The window caps what one request can see, but you pay for every token in it on every call, and long prompts take longer to start answering. Treat a big window as room, not a target: retrieve only the relevant passages with RAG, or cache a repeated prefix, rather than resending everything.

More context also isn’t free for quality. Anthropic’s docs say accuracy and recall degrade as the token count grows, which it calls “context rot”, so what you put in the window matters as much as how much fits.

Don’t mix up

Common confusions

Context window vs max output tokens
The window is the whole budget for one request. Max output tokens is a separate, much smaller cap on the reply. A model with a million-token window still stops when the reply reaches its output cap.
Context window vs what the model knows
What a model learned in training is not in the window, and nothing in the window is remembered after the request ends unless your app sends it again. Anthropic describes the window as the model’s “working memory”.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary