Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Models and context

What is chain-of-thought prompting?

Also called: chain-of-thought prompting, CoT, step-by-step prompting

Definition

Chain-of-thought prompting is asking a language model to write out intermediate reasoning steps before its final answer, usually by showing worked examples or telling it to think step by step.

Explained

How it works

The technique was named in a 2022 Google paper, “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models” (NeurIPS 2022). Its few-shot examples showed the working, not just the answer. With eight such examples, a 540-billion-parameter model reached state-of-the-art accuracy on GSM8K, a benchmark of maths word problems, and the gains appeared only in sufficiently large models.

Reasoning models now do this on their own: they generate reasoning tokens before answering, billed as output and mostly hidden. OpenAI says prompting them to “think step by step” is unnecessary. Anthropic treats manual chain-of-thought as a fallback for when thinking is off, and warns that on Claude Fable 5.1, Fable 5, Opus 5.5, Opus 5 and Sonnet 5.5 a prompt asking the model to write out its reasoning may be declined.

Example

The apples question from the paper

The paper’s first figure asks: “The cafeteria had 23 apples. If they used 20 to make lunch and bought 6 more, how many apples do they have?” With a standard example in the prompt, the model answered 27, which is wrong. With an example that showed its working, it wrote out the steps and reached 9.

The working costs tokens. With o200k_base, the short answer is 8 tokens and the worked one 55, about 7 times as many. Output is the slow, expensive side of a request: OpenAI’s latency guide says cutting half your output tokens may cut about half your latency.

The two answers (8 and 55 tokens, o200k_base)
Standard prompting:
A: The answer is 27.

Chain-of-thought prompting:
A: The cafeteria had 23 apples originally. They used 20 to make lunch. So they had 23 - 20 = 3. They bought 6 more apples, so they have 3 + 6 = 9. The answer is 9.

Cost and quality

Why it matters

On models without built-in reasoning, asking for the working can still improve multi-step answers, and the steps make mistakes easier to spot. On reasoning models, spend the effort on the reasoning setting instead (OpenAI’s reasoning.effort, Anthropic’s effort and thinking options) and keep prompts plain.

Either way you pay for every reasoning token, so use it where a task has real steps, such as maths, planning or debugging, not for lookups or classification.

Don’t mix up

Common confusions

Chain-of-thought prompting vs reasoning models
Chain-of-thought is a prompt technique: the steps appear in the visible answer. A reasoning model is trained to think before answering, whatever the prompt says, and its reasoning tokens are usually hidden but still billed.
Chain-of-thought vs few-shot prompting
Chain-of-thought can be few-shot (examples that include the working, as in the original paper) or a plain instruction to reason first. Plain few-shot prompting only shows inputs and answers.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary