AI glossary · Models and context
What are max output tokens?
Also called: max_tokens, max_output_tokens, max_completion_tokens, maxOutputTokens, output token limit
Definition
Max output tokens is the cap on how many tokens a model may generate in one response, set per request up to the model’s own limit, with any reasoning tokens counted inside it.
Explained
How it works
Every API has the setting under a different name: max_output_tokens in OpenAI’s Responses API, max_completion_tokens in Chat Completions, max_tokens in Anthropic’s Messages API (where it is required) and maxOutputTokens in Gemini. You choose a value up to the model’s maximum. It is a ceiling, not a target: the model normally stops sooner, when it has finished.
On reasoning models the cap covers the thinking too. OpenAI’s reference says max_output_tokens includes reasoning tokens, and Anthropic says thinking tokens count toward max_tokens. Set it too low and the budget can run out before any visible answer appears, while you still pay for the input and the thinking.
A reply that hits the cap is cut off mid-sentence, and the API says so: finish_reason: "length" in Chat Completions, status: "incomplete" with incomplete_details.reason: "max_output_tokens" in Responses, stop_reason: "max_tokens" from Anthropic and finishReason: "MAX_TOKENS" from Gemini. Check it before you parse a reply, especially JSON.
Example
Output caps next to context windows
In our daily data the output cap is a small slice of the window. GPT-6.1 Sol reads up to 1,050,000 tokens but writes at most 128,000 in one reply (12.2% of its window). That is still about 111,000 words of English prose at our measured o200k_base ratio, far more than most answers need. Gemini 3.8 Flash stops at 65,536.
Some providers publish no separate cap, as with Grok 4.7; the reply then only has to fit in what is left of the context window.
| Model | Context window | Max output | Share of window |
|---|---|---|---|
| Claude Sonnet 5.5 | 1,000,000 | 128,000 | 12.8% |
| GPT-6.1 Sol | 1,050,000 | 128,000 | 12.2% |
| Gemini 3.8 Flash | 1,048,576 | 65,536 | 6.3% |
| Llama 4 Maverick | 128,000 | 16,384 | 12.8% |
| Grok 4.7 | 500,000 | No separate cap published | n/a |
From our daily data, 2026-10-11. The model comparison table lists every model.
Cost and quality
Why it matters
Output usually costs several times as much as input per token (on Claude Sonnet 5.5, $10 against $2 per million), and it is the slow part of a request. A sensible cap protects your bill and latency from runaway replies.
Too tight a cap causes its own failures: truncated answers, half-written code and JSON that won’t parse, which look like model mistakes until you read the stop reason. Leave generous room on reasoning models; OpenAI suggests reserving at least 25,000 tokens for reasoning and output when you start.
Don’t mix up
Common confusions
- Max output tokens vs context window
- The context window is the total for input plus output. The output cap limits only the reply, and the reply must also fit in whatever room the input leaves.
- Max output tokens vs asking for a short answer
- “Answer in under 100 words” shapes the reply so it ends cleanly. The token cap just cuts it off. Use the instruction to get short answers and the cap as a safety limit.
Go deeper
Try it and read more
- Free toolAI Model ComparisonCompare prices, context windows and features across models.
- Free toolGemini API PlaygroundTry the Gemini API free with your own key.
- Free toolJSON RepairFix broken JSON from LLM output.
- Guide · 10 min readContext windows explained: what counts, what happens at the limit, and how to check fitWhat an LLM context window is, what counts toward it, max output vs context, what happens when you exceed it, and when RAG beats long context.
- Guide · 10 min readHow to fix invalid JSON from LLMs (and stop getting it)Why ChatGPT, Claude and Gemini return broken JSON, and how to fix it: structured outputs, schema validation, retries, safe repair and streaming.
Related
Related terms
- Context windowA context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.
- Reasoning tokensReasoning tokens are the tokens a reasoning model generates while it works through a problem before writing its answer, billed as output tokens even though the API hides them or returns only a summary.
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- Structured outputsStructured outputs is an LLM API feature that constrains the model’s reply to a JSON Schema you supply, so the response parses and has the fields and types you asked for, unlike JSON mode, which only promises valid JSON.