Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Models and context

What are max output tokens?

Also called: max_tokens, max_output_tokens, max_completion_tokens, maxOutputTokens, output token limit

Definition

Max output tokens is the cap on how many tokens a model may generate in one response, set per request up to the model’s own limit, with any reasoning tokens counted inside it.

Explained

How it works

Every API has the setting under a different name: max_output_tokens in OpenAI’s Responses API, max_completion_tokens in Chat Completions, max_tokens in Anthropic’s Messages API (where it is required) and maxOutputTokens in Gemini. You choose a value up to the model’s maximum. It is a ceiling, not a target: the model normally stops sooner, when it has finished.

On reasoning models the cap covers the thinking too. OpenAI’s reference says max_output_tokens includes reasoning tokens, and Anthropic says thinking tokens count toward max_tokens. Set it too low and the budget can run out before any visible answer appears, while you still pay for the input and the thinking.

A reply that hits the cap is cut off mid-sentence, and the API says so: finish_reason: "length" in Chat Completions, status: "incomplete" with incomplete_details.reason: "max_output_tokens" in Responses, stop_reason: "max_tokens" from Anthropic and finishReason: "MAX_TOKENS" from Gemini. Check it before you parse a reply, especially JSON.

Example

Output caps next to context windows

In our daily data the output cap is a small slice of the window. GPT-6.1 Sol reads up to 1,050,000 tokens but writes at most 128,000 in one reply (12.2% of its window). That is still about 111,000 words of English prose at our measured o200k_base ratio, far more than most answers need. Gemini 3.8 Flash stops at 65,536.

Some providers publish no separate cap, as with Grok 4.7; the reply then only has to fit in what is left of the context window.

Context window and maximum output per reply
ModelContext windowMax outputShare of window
Claude Sonnet 5.51,000,000128,00012.8%
GPT-6.1 Sol1,050,000128,00012.2%
Gemini 3.8 Flash1,048,57665,5366.3%
Llama 4 Maverick128,00016,38412.8%
Grok 4.7500,000No separate cap publishedn/a

From our daily data, 2026-10-11. The model comparison table lists every model.

Cost and quality

Why it matters

Output usually costs several times as much as input per token (on Claude Sonnet 5.5, $10 against $2 per million), and it is the slow part of a request. A sensible cap protects your bill and latency from runaway replies.

Too tight a cap causes its own failures: truncated answers, half-written code and JSON that won’t parse, which look like model mistakes until you read the stop reason. Leave generous room on reasoning models; OpenAI suggests reserving at least 25,000 tokens for reasoning and output when you start.

Don’t mix up

Common confusions

Max output tokens vs context window
The context window is the total for input plus output. The output cap limits only the reply, and the reply must also fit in whatever room the input leaves.
Max output tokens vs asking for a short answer
“Answer in under 100 words” shapes the reply so it ends cleanly. The token cap just cuts it off. Use the instruction to get short answers and the cap as a safety limit.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary