Tokens & Costs
Context window checker: will your text fit?
Paste text or enter a token count to see whether it fits each model’s context window, with room left for the reply. Counting runs in your browser.
0 characters · 0 words
Counting runs in your browser. Your text isn’t sent anywhere unless you ask a provider for an official count.
Tokens the model may write. Reasoning models spend thinking tokens here too.
System prompt, tool definitions, chat history.
Will it fit?
Other promptYour textUncertainReply
GPT-6.1 Sol
OpenAI · 1,050,000-token window
Add some text to check.
GPT-6 Luna
OpenAI · 1,050,000-token window
Add some text to check.
Claude Opus 5.5
Anthropic · 1,000,000-token window
Add some text to check.
Free, uses your own API key. Sends this text to Anthropic.Claude Sonnet 5.5
Anthropic · 1,000,000-token window
Add some text to check.
Free, uses your own API key. Sends this text to Anthropic.Claude Haiku 5.5
Anthropic · 1,000,000-token window
Add some text to check.
Free, uses your own API key. Sends this text to Anthropic.Gemini 3.8 Flash
Google · 1,048,576-token window
Add some text to check.
Free, uses your own API key. Sends this text to Google.DeepSeek V4.1 Flash
DeepSeek · 1,048,576-token window
Add some text to check.
Llama 4 Maverick
Meta · 128,000-token window
Add some text to check.
Steps
How to use the context window checker
- Paste your text or drop a file, or switch to “Enter a token count” if you already know it.
- Set the room for the reply, and any tokens already used by a system prompt, tool definitions or chat history.
- Read the verdict for each model: “Fits”, “Too close to call” or “Too large”, with how much space is left.
- For Claude and Gemini, get the official count with your own API key to turn a range into an exact answer.
- If it doesn’t fit, add one of the suggested larger-context models, or split the text into the number of parts shown.
Method
How it works
A model’s context window is the most tokens it can handle in one request, including its own reply. Windows range from a few thousand tokens to 2,000,000 for Grok 4.20, the largest in our data. The checker adds up everything a request carries and compares it with each model’s published limit.
What has to fit
A request is more than the text you paste. It also carries the system prompt, any tool or function definitions, the earlier messages in a conversation, and room for the answer. The checker splits these into three parts: your text, other prompt tokens, and the reply. The bar for each model shows all three against its window, and a line marks the edge when the request goes past it.
How the tokens are counted
Different models split the same text into different numbers of tokens, so the checker counts per model. For OpenAI models and for open-weight models whose tokenizers are published (such as DeepSeek, Qwen and Mistral), it runs the model’s own tokenizer in your browser, the same engine as the AI token counter, and the count is exact.
Claude, Gemini and many other models don’t publish their tokenizers. Rather than pretend to know their count, the checker gives a range. We measured the published modern tokenizers against OpenAI’s on more than 1,500 texts in many languages: they used between 0.61 and 2.0 times as many tokens, and most stayed within 0.88 to 1.43 times. So the range runs from 0.5× to 2× OpenAI’s count, widened to 3× when most of the letters are in a non-Latin script, where tokenizers differ most. A model is only marked “Fits” if the top of the range fits, and only “Too large” if the bottom of the range doesn’t. In between, it says “Too close to call”, and the official count (free, with your own API key) gives the exact answer.
If it doesn’t fit
You have three main options. Choose a model with a larger window: the checker lists the cheapest current models with room for your request. Split the text into parts that each fit, and process them one at a time; the checker tells you how many parts a given model needs. Or send less: summarise long documents first, or retrieve only the relevant passages instead of the whole text. The last option is often the best, because very long prompts cost more, run slower, and models use information in the middle of long inputs less reliably.
Context windows and reply limits come from the OpenRouter models API and the LiteLLM model list, refreshed daily (last updated 2026-10-09). Some providers offer larger windows on request or on special tiers; the checker uses the standard published limit.
Examples
Worked examples
How much text fits in popular models
| Model | Context window | Longest reply | About (English) |
|---|---|---|---|
| GPT-6.1 Sol OpenAI | 1,050,000 | 128,000 | 787,500 words · 1,575 pages |
| GPT-6 Luna OpenAI | 1,050,000 | 128,000 | 787,500 words · 1,575 pages |
| Claude Opus 5.5 Anthropic | 1,000,000 | 128,000 | 750,000 words · 1,500 pages |
| Claude Sonnet 5.5 Anthropic | 1,000,000 | 128,000 | 750,000 words · 1,500 pages |
| Claude Haiku 5.5 Anthropic | 1,000,000 | 128,000 | 750,000 words · 1,500 pages |
| Gemini 3.8 Flash Google | 1,048,576 | 65,536 | 786,432 words · 1,573 pages |
| DeepSeek V4.1 Flash DeepSeek | 1,048,576 | – | 786,432 words · 1,573 pages |
| Llama 4 Maverick Meta | 128,000 | 16,384 | 96,000 words · 192 pages |
Words and pages use ¾ of a word per token and 500 words per page, before leaving room for the reply. “–” means no published reply limit. Data updated 2026-10-09.
FAQ
Frequently asked questions
What is a context window?
It is the most text a model can handle in one request, measured in tokens. Everything counts: the system prompt, tool definitions, the conversation so far, any documents you attach, and the reply the model writes.
Does the reply count towards the context window?
Yes, for most providers the prompt and the reply have to fit in the window together, so a model with a 200,000-token window can’t read 199,000 tokens and then write 4,000. This checker always leaves the room you set for the reply. Each model also has a separate limit on how long one reply can be, and the checker warns you if your reply room is above it.
How many pages fit in 200,000 tokens?
For English prose, roughly 150,000 words or about 300 printed pages, using the common estimate of three-quarters of a word per token and 500 words per page. Code, tables and other languages use more tokens per word, so check your own text.
Why does Claude or Gemini show a range instead of a number?
Their tokenizers aren’t public, so the exact count isn’t knowable without asking the provider. The checker takes OpenAI’s count of your text and allows from 0.5× to 2× of it (3× for text mostly in non-Latin scripts), the spread we measured across the tokenizers that are public. It only says “Fits” if even the top of the range fits. Click “Get the official count” with your own API key for an exact answer.
What happens if a request is too long?
Most APIs reject it with an error, such as Anthropic’s “prompt is too long”. Chat apps usually handle it for you by dropping or summarising older messages, which means the model quietly stops seeing part of the conversation.
Do reasoning or thinking tokens use the context window?
Yes. Reasoning models write hidden thinking tokens before the visible answer, and those count as output in the same window. If you use a reasoning model, set the reply room high enough for the thinking as well as the answer.
Is a bigger context window always better?
Not always. Long prompts cost more (some models charge a higher rate above a size threshold), take longer to process, and models tend to use information buried in the middle of very long inputs less reliably than information at the start or end. Sending only the relevant parts often gives better answers for less money.
Related
Related tools
- AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- AI Model ComparisonCompare prices, context windows and features across models.
- AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.
- Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.
- GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.