Guide · Coding agents
How to reduce Claude Code token usage, and what each fix costs you
Claude Code resends the whole conversation, plus its system prompt, tools and CLAUDE.md, with every step, so token use grows with context size more than with what you type. The biggest savings come from clearing between tasks, using Sonnet or a lower effort for routine work, keeping CLAUDE.md and MCP servers lean, and asking specific questions that avoid broad file reads.
By Tahir NazirUpdated 11 min read
On this page
- Why does Claude Code use so many tokens?
- How do I check token usage in Claude Code?
- Should I use /clear or /compact?
- Which model and effort level use the fewest tokens?
- Keep CLAUDE.md, MCP servers and skills lean
- Write prompts that don’t trigger broad file reads
- Use subagents for noisy work, on a smaller model
- What prompt caching already does for you
- Questions people ask
Why does Claude Code use so many tokens?
Because the model remembers nothing between requests, Claude Code sends the full context every time: the system prompt, tool definitions, your CLAUDE.md, every earlier message and tool result, and your new message. And it sends a request at every step of its loop, each time it reads a file, runs a command or replies. A one-line question late in a long session still carries everything before it.
1 · System prompt
Changes when the set of loaded tools changes
- Core instructions and built-in toolsEvery requestSet by Claude Code
- MCP tool names and server instructionsEvery request; full definitions only when a tool is used/mcp disable
2 · Project context
Changes at session start, /clear or /compact
- CLAUDE.md files and rules without pathsSession start, in fullUnder 200 lines
- Auto memory (MEMORY.md)First 200 lines or 25KB/memory
3 · Conversation
Grows every turn
- Skill descriptionsListed at session start, one line each; the full skill only when usedFewer skills
- Your prompts and Claude’s repliesEverything since the last /clear or /compact/clear, /compact
- Tool results: files read, command outputAppended each time a tool runsSpecific prompts, subagents
- ResponseClaude’s reply and its thinking, billed as output/effort, /model
Layers 1 to 3 are re-sent with every request; unchanged parts are read from the prompt cache
Prompt caching softens this. The unchanged start of each request is read from a cache, and on Opus 5.5 a cached token costs $0.20 per million instead of $4. But cheap isn’t free. In one session we measured on 2026-10-08, 940 requests over 8.4 active hours, never cleared, the median request carried 426,571 tokens and the largest 966,324. Cache reads were 99.5% of all input tokens and about 72% of what the session would cost at API list prices.
How do I check token usage in Claude Code?
Two commands tell you most of what you need. Run them before you change anything, so you know which fix matters for you.
- `/context` draws the current context window as a coloured grid, with a breakdown by category (including memory files, MCP tools and skills) and suggestions for heavy tools and bloated memory. In fullscreen mode,
/context allshows the full breakdown. - `/usage` (also
/costand/stats) shows the session’s tokens by model, an estimated cost at list price, and a prompt cache line with the share of input served from the cache and any misses. On Pro, Max, Team and Enterprise it adds plan usage bars and a breakdown by skill, subagent, plugin and MCP server. - The status line can show context use all the time: Claude Code passes
context_window.used_percentageto your status line script.
The dollar figure in /usage is computed on your machine, so treat it as an estimate. On the API, the Console’s usage page is the real bill.
Should I use /clear or /compact?
Use /clear when you switch to unrelated work, and /compact when you need to keep going on the same task with less history. Anthropic’s cost docs name clearing between unrelated tasks as one of the two highest-impact habits, because stale context costs tokens on every later message.
- `/clear` starts a new conversation with empty context. It costs nothing. Run
/renamefirst if you want to find the old session again with/resume. - `/compact` replaces the history with a summary. Add focus instructions, such as
/compact Focus on code samples and API usage, to choose what survives. The summary is a request that reads the whole conversation: cheap while the cache is warm, expensive after a break. - `/rewind` (or Escape twice) goes back to an earlier checkpoint when Claude went down the wrong path. It returns to a prefix that is already cached, so it is cheaper than compacting.
You can also tell compaction what to keep in your project’s CLAUDE.md, as the docs suggest:
# Compact instructions
When you are using compact, please focus on test output and code changesWhat does clearing save? Our measured session’s cache reads alone come to $88.02 at Opus 5.5 list prices. If clearing between tasks had held each request to 100,000 tokens, the same 940 requests would have read about $18.80. That’s a rough model, not a measurement, but the direction is clear. Claude Code also compacts automatically as the window fills; on 1M-context models, /autocompact 200k sets a lower threshold, trading smaller requests for more frequent summaries. If you see Prompt is too long, the window filled anyway.
Which model and effort level use the fewest tokens?
The model sets the price of every token, and the effort level sets how many thinking tokens Claude spends. Anthropic’s docs say Sonnet handles most coding tasks well and costs less than Opus, and suggest reserving Opus for complex architectural decisions or multi-step reasoning. Claude Code’s default on Pro, Max, Team, Enterprise and the API is Opus 5.5.
| Model | Input | Cached input | Output |
|---|---|---|---|
| Claude Opus 5.5 | $4 | $0.20 | $20 |
| Claude Sonnet 5.5 | $2 | $0.10 | $10 |
| Claude Haiku 5.5 | $0.10 | $0.01 | $0.50 |
List prices per million tokens as of 2026-10-09. Haiku 5.5 charges more per token once a prompt passes 100,000 tokens. On a subscription, the model and effort you choose also affect how fast you use your limits.
- Switch with `/model`, for example
/model sonnet. Opus 5.5 output costs twice as much as Sonnet 5.5’s. Choose at the start of a session: each model has its own cache, so a mid-session switch re-reads the whole conversation uncached. - Lower the effort with
/effort lowor/effort mediumfor routine edits. Thinking is billed as output. Opus 5.5, Sonnet 5.5 and Haiku 5.5 default tomedium, and their thinking can’t be turned off. - Try `opusplan`, which uses Opus in plan mode and Sonnet for the edits. Each switch between them is a model change, so it starts a fresh cache.
On models with a fixed thinking budget, the MAX_THINKING_TOKENS environment variable caps it; adaptive models ignore it, so use effort there. With an API key or a subscription, changing effort on the 5.5 models keeps the cache; on most other models and on Bedrock or Google Cloud it doesn’t.
Keep CLAUDE.md, MCP servers and skills lean
Everything that loads at session start rides along on every request until you clear. Trim it once and every step gets cheaper.
- CLAUDE.md under 200 lines. It loads in full.
@pathimports don’t reduce the cost, since imported files load at launch too. Move instructions for one part of the codebase into.claude/rules/files withpaths:frontmatter, which load only when Claude touches matching files, and move occasional workflows (PR reviews, migrations) into skills, which load only when used. Our CLAUDE.md generator warns past 200 lines. - Disable MCP servers you aren’t using with
/mcp disable <name>. By default only tool names and server instructions load until a tool is used, but that deferral (tool search) switches off whenANTHROPIC_BASE_URLpoints to a third-party gateway, and then every definition loads up front. - Prefer command-line tools such as
gh,awsorgcloudwhen they exist: the docs call them more context-efficient than MCP servers because they add no per-tool listing. - Leave notes for humans in HTML comments. Block-level
<!-- -->comments in CLAUDE.md are stripped before Claude sees them. - Run `/doctor` now and then: its checkup weighs unused skills, MCP servers and plugins against their context cost, and proposes cuts to CLAUDE.md content Claude could work out from the code.
Edits to CLAUDE.md take effect after the next /clear, /compact or restart. To see what a file costs, paste it into the token counter; the context window checker shows how much of a model’s window a document takes before you add it to a session.
Write prompts that don’t trigger broad file reads
Tool results, mostly file contents and command output, are usually the part of the context that grows fastest. The docs’ example: “improve this codebase” triggers broad scanning, while “add input validation to the login function in auth.ts” lets Claude work with minimal file reads.
- Plan first on big tasks. Press Shift+Tab to reach plan mode, so Claude proposes an approach before reading and editing widely.
- Stop early. Press Escape as soon as Claude heads the wrong way; every wasted step adds to the context.
- Give a target. Tests or expected output let Claude check its own work instead of you asking for fixes.
- Install a code intelligence plugin for typed languages: one “go to definition” call replaces a search plus several file reads.
To keep Claude out of build output or generated files, don’t rely on a .claudeignore file: the permissions docs say it has no effect. Add Read deny rules to .claude/settings.json instead. Claude applies them to its file tools on a best-effort basis, including Grep and Glob:
{
"permissions": {
"deny": [
"Read(./dist/**)",
"Read(./coverage/**)",
"Read(./.env)"
]
}
}Use subagents for noisy work, on a smaller model
Running tests, fetching documentation or reading logs can flood the main conversation. A subagent does that work in its own context and returns only a summary, so your main context stays small. Ask for it directly, as in the docs’ example (“Use a subagent to run the test suite and report only the failing tests with their error messages”), or define one with a cheaper model:
---
name: test-runner
description: Runs the test suite and reports only failing tests with their errors
tools: Bash, Read, Grep
model: haiku
---
Run the tests, then reply with each failing test and its error message. Nothing else.The trade-off: a subagent’s requests still count, and it starts with an empty cache of its own. Subagents get a five-minute cache lifetime even on a subscription. Agent teams cost more again: Anthropic says they use about 7x the tokens of a standard session when teammates run in plan mode.
What prompt caching already does for you
Claude Code turns on prompt caching automatically and orders each request so the parts that rarely change come first: system prompt and tools, then project context, then the conversation. A change early in that order invalidates everything after it, which is why these actions cost one slower, full-price step:
- switching model, or changing effort on most models other than the 5.5 generation
- turning on fast mode, compacting, or upgrading Claude Code
- connecting or removing an MCP server, when tool definitions load up front
The cache also expires when you stop. On a subscription within your plan, the main conversation’s cache lasts an hour; on an API key, a cloud provider or usage credits it lasts five minutes. API users can set promptCacheTtl to 1h, which costs more per cache write and pays off if you often pause for longer than five minutes. Our prompt caching guide explains the pricing.
FAQ
Questions people ask
How do I see how many tokens Claude Code is using?
Run /context for a breakdown of what fills the current context window, and /usage (or its alias /cost) for the session’s tokens by model, an estimated cost and prompt cache hits. On a Claude plan, /usage also shows your session and weekly limit bars. For API billing, the Console usage page is authoritative.
Does /clear delete my code or my conversation?
No. /clear only empties the context Claude sees; your files are untouched. The previous conversation is saved, so you can return to it with /resume. Run /rename first to give it a name that’s easy to find. Clearing costs no tokens, which makes it the cheapest fix for a bloated session.
Is /compact or /clear better for saving tokens?
/clear saves more: it costs nothing and starts from an empty context. /compact keeps a summary so you can continue the same task, but producing that summary is a request over the whole conversation. Use /compact with focus instructions at a natural break in a task, and /clear when you switch to unrelated work.
Do MCP servers use tokens if I don’t call them?
A little. By default Claude Code defers MCP tool definitions, so only tool names and server instructions load until Claude uses a tool. If tool search is off, for example behind a third-party gateway set with ANTHROPIC_BASE_URL, every definition loads into every request. Disable unused servers with /mcp disable <name>.
Does a .claudeignore file stop Claude Code reading files?
No. Claude Code’s permissions documentation says a .claudeignore file has no effect. To block reads, add Read deny rules such as Read(./dist/**) to the permissions.deny list in .claude/settings.json. Claude applies them to its file tools, including Grep and Glob, on a best-effort basis.
Does switching models mid-session cost more?
Yes, once. Each model has its own prompt cache, so the first request after /model re-reads the whole conversation with no cache hits. While the cache is warm, Claude Code asks you to confirm the switch. Pick the model at the start of a session, or after /clear, when there is little context to re-read.
Try it
Tools from this guide
Keep reading