AI glossary · Tokens and cost
What are reasoning tokens?
Also called: thinking tokens, hidden reasoning
Definition
Reasoning tokens are the tokens a reasoning model generates while it works through a problem before writing its answer, billed as output tokens even though the API hides them or returns only a summary.
Explained
How it works
Reasoning models, such as OpenAI’s GPT-5 models and Claude with thinking on, write a private chain of reasoning before the reply. That reasoning is made of ordinary output tokens. OpenAI never returns the raw reasoning but can send a summary if you ask for one, and Claude can return a summary or nothing at all. Either way, you pay for every token the model actually generated, not for what you see.
The count is in each response’s usage data: output_tokens_details.reasoning_tokens in OpenAI’s Responses API and output_tokens_details.thinking_tokens in Claude’s Messages API, both already included in the output total. OpenAI says a request can use anywhere from a few hundred to tens of thousands of them.
You steer the amount with an effort setting (reasoning.effort on OpenAI, output_config.effort on Claude). Reasoning also counts against the output cap: if max output tokens is too low, the model can run out while still thinking, before writing any visible answer. OpenAI suggests reserving at least 25,000 tokens for reasoning and output when you start.
Example
6,000 hidden tokens behind a 500-token answer
Suppose a request to GPT-6.1 Sol sends 2,000 input tokens, gets a 500-token answer, and the usage data reports 6,000 reasoning tokens. The token counts are illustrative; the prices are from our daily data: $2 input and $10 output per million tokens.
Billed output is 6,500 tokens, not 500, so the request costs $0.069 instead of the $0.009 the visible reply suggests: 7.7 times as much. At 1,000 requests a day, that gap is $60.00 a day.
| Tokens | Cost | |
|---|---|---|
| Input | 2,000 | $0.004 |
| Reasoning (hidden, billed as output) | 6,000 | $0.06 |
| Visible answer | 500 | $0.005 |
| Total | 8,500 | $0.069 |
Prices from our daily data, 2026-10-11. The reasoning count varies by request and effort setting.
Cost and quality
Why it matters
Reasoning tokens are often most of a reasoning model’s bill, and they make cost per request hard to predict, because the model decides how long to think. An estimate based only on the visible answer can be several times too low.
Use lower effort for simple tasks, set an output cap with room for thinking, and log the reasoning count from the usage data so you can see where the money goes. The LLM cost calculator prices a workload once you include thinking in the output tokens.
Don’t mix up
Common confusions
- Reasoning tokens vs chain-of-thought prompting
- Chain-of-thought prompting asks a model to show its steps in the visible answer. Reasoning tokens come from a model trained to think before answering; they are usually hidden and billed whether you see them or not.
- Billed thinking vs the summary you see
- Anthropic bills the full thinking Claude generated, not the summary it returns, so the billed output count won’t match the visible text. Read
thinking_tokensin the usage data for the real number.
Go deeper
Try it and read more
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Free toolAI Model ComparisonCompare prices, context windows and features across models.
- Guide · 11 min readClaude vs GPT vs Gemini pricing, explained with live pricesClaude, GPT and Gemini API prices side by side from daily data: how each bills caching, batch, long prompts and thinking, with a worked cost example.
- Guide · 11 min readHow to estimate LLM API costs: the formula, worked examples and the trapsEstimate LLM API costs per request and per month: the token formula, worked examples for chatbots, RAG and agents, plus caching, batch and reasoning tokens.
Related
Related terms
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- Chain of thoughtChain-of-thought prompting is asking a language model to write out intermediate reasoning steps before its final answer, usually by showing worked examples or telling it to think step by step.
- Max output tokensMax output tokens is the cap on how many tokens a model may generate in one response, set per request up to the model’s own limit, with any reasoning tokens counted inside it.
- Context windowA context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.