AI glossary · Agents and tools
What is an AI agent?
Also called: LLM agent, agentic loop, agent loop
Definition
An AI agent is a program in which a language model works towards a goal by repeatedly choosing a tool to call, reading the result and deciding the next step, until the task is done or a stop condition is reached.
Explained
How it works
Anthropic draws a useful line between workflows, where your code fixes the sequence of model and tool calls in advance, and agents, where the model decides its own steps and which tools to use. Underneath, an agent is a loop: send the goal and the available tools, let the model ask for a tool call, run it, send back the result, and repeat.
The model never runs anything itself. Your code, often called the harness, executes each call and returns real results from the environment, such as test output or a file’s contents. That feedback is what lets an agent notice a mistake and fix it. The loop ends when the model replies without asking for a tool, or when a limit you set, such as a maximum number of turns, stops it.
Coding agents such as Claude Code are this loop with file, search, shell and web tools, plus context management: project instructions in CLAUDE.md, subagents for side tasks, and compaction when the context window fills up.
Example
Token growth in a 7-request coding loop
Claude Code’s docs walk through “fix the failing tests” as six tool calls: run the tests, read the error, search for the source, read it, edit it, and run the tests again. Each call is a new request, and each request sends the original messages, every earlier reply and every tool result again.
Assume a 15,000-token start (system prompt, tool definitions, project instructions and your message) and about 2,000 tokens added per tool round trip. These sizes are illustrative; the shape is not. The last request carries 27,000 tokens, but the seven together send 147,000, which is 5.4 times as much. At Claude Sonnet 5.5’s input price of $2 per million tokens that is $0.29 of input for one small fix, before output and before any prompt caching discount.
| Request | What the agent does | Input tokens sent |
|---|---|---|
| 1 | Run the test suite | 15,000 |
| 2 | Read the failure output | 17,000 |
| 3 | Search for the relevant source files | 19,000 |
| 4 | Read those files | 21,000 |
| 5 | Edit the code | 23,000 |
| 6 | Run the tests again | 25,000 |
| 7 | Reply with the fix (no tool call) | 27,000 |
| Total | 147,000 |
Price from our daily data, 2026-10-11. Output tokens are not included.
Cost and quality
Why it matters
Because every turn re-sends the conversation, an agent’s input bill grows roughly with the square of its number of turns when each turn adds a similar amount. Long tasks, large tool results and many tool definitions all multiply it. Prompt caching, shorter tool outputs and fewer turns are the main levers.
Quality depends on the feedback the agent can get. Anthropic stresses that agents need ground truth from the environment at each step, such as tests that pass or fail, and recommends starting with the simplest solution and adding an agent only when a fixed workflow falls short.
Don’t mix up
Common confusions
- AI agent vs chatbot
- A chatbot answers from what it was given in one reply. An agent takes actions through tools and loops on the results, so it can run code, edit files or query systems before it answers.
- Agent vs workflow
- In a workflow, your code decides the order of steps and the model fills in each one. In an agent, the model decides the steps. Workflows are more predictable and cheaper; agents handle tasks whose steps you can’t list in advance.
Go deeper
Try it and read more
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Free toolPrompt Caching CalculatorEstimate savings from prompt caching.
- Free toolContext Window CheckerSee whether your text fits each model’s context window.
- Guide · 10 min readHow to choose a model for coding agents: the criteria that matterFind the best model for a coding agent by tool-use reliability, context, speed, price per task and where it runs, with live prices for tool-capable models.
- Guide · 11 min readHow to reduce Claude Code token usage, and what each fix costs youCut Claude Code token use with /clear, /compact, model and effort choice, a lean CLAUDE.md, fewer MCP servers and subagents, and the trade-off of each.
Related
Related terms
- Function callingFunction calling, also called tool use, is an LLM API feature in which the model replies with a structured request to run a function you described, with JSON arguments, which your code executes before sending the result back.
- SubagentA subagent is a separate AI worker that a main agent hands a task to, running in its own context window with its own system prompt, tools and model, and returning only its result to the main conversation.
- MCPMCP stands for Model Context Protocol, an open standard that lets AI applications such as Claude Code, ChatGPT or VS Code connect to outside tools and data through one common interface.
- Context windowA context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.
- Prompt cachingPrompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.