Guide · Coding agents
How to choose a model for coding agents: the criteria that matter
Choose a coding-agent model on the job it has to do, not on a leaderboard: reliable tool calls, enough context and output room, acceptable speed, and the lowest price per finished task. Agents resend their whole context on every step, so input tokens and prompt caching decide the bill far more than the output price does.
By Tahir NazirUpdated 10 min read
On this page
- What matters when choosing a model for a coding agent?
- Why agents cost more than chat, and how to price a task
- Tool-use reliability
- Context window and output limit
- Speed and reasoning modes
- Where you can run it: APIs, clouds and open weights
- What benchmarks like SWE-bench tell you
- A simple way to choose
- Questions people ask
What matters when choosing a model for a coding agent?
A coding agent is a loop: the model reads files, calls tools, edits code and runs tests until the task is done. The model you want is the cheapest one that finishes your tasks reliably inside that loop. Seven criteria decide it:
- Tool-use reliability: well-formed calls, the right tool, no loops.
- Context window: room for the repository slices, tool output and history a task accumulates.
- Output limit: room to write a large file or a long patch in one reply.
- Speed: how long a task takes end to end, which shapes how you work with the agent.
- Price per task: what a finished task costs, including caching, not the price per token.
- Reasoning controls: whether you can turn thinking up for hard steps and down for easy ones.
- Where it runs: the provider’s API, a cloud you already use, or open weights you host yourself.
Rankings change often, so this guide doesn’t name a winner. It shows how to judge each criterion and gives live prices for the current tool-capable models.
Why agents cost more than chat, and how to price a task
Every step of the loop is a new API request that resends everything so far: the system prompt, tool definitions, the task, every file read and every test log. Here is one small example task, a bug fix in eight steps:
- 1Search the code12,500 in · 600 out
- 2Read two files16,100 in · 300 out
- 3Read a test file25,400 in · 200 out
- 4Edit the code29,600 in · 2,500 out
- 5Run the tests32,400 in · 150 out
- 6Fix the failure38,550 in · 1,200 out
- 7Run the tests again40,050 in · 150 out
- 8Summarise the change42,700 in · 600 out
- Input billed
- 237,300
- Output billed
- 5,700
- Final context
- 42,700
Example task · tokens per request · output includes thinking
The task bills 237,300 input tokens against 5,700 output tokens. Without caching, input is 89% of the cost on Claude Sonnet 5.5, even though its output price is 5× its input price. With caching, each step pays the cheap cached rate for everything it re-sent, and the bill for this task falls by 23% to 73% across the models below that list a cache price. A longer task has more steps, each resending more, so the gap grows.
| Model | Open weights | Context window | Input / cached / output per 1M | Task, no caching | Task, cached |
|---|---|---|---|---|---|
| gpt-oss-120b | Yes | 131,072 | $0.037 / – / $0.17 | $0.00975 | $0.00975 |
| Claude Haiku 5.5 | No | 1,000,000 | $0.10 / $0.01 / $0.50 | $0.0266 | $0.0101 |
| GPT-6 Luna | No | 1,050,000 | $0.10 / $0.01 / $0.50 | $0.0266 | $0.0101 |
| Qwen3.8 Flash | Yes | 1,000,000 | $0.15 / $0.016 / $0.47 | $0.0383 | $0.0143 |
| DeepSeek V4.1 Flash | Yes | 1,048,576 | $0.30 / $0.006 / $1.20 | $0.078 | $0.0208 |
| Mistral Large 4 | No | 1,048,576 | $0.68 / $0.07 / $2.09 | $0.17 | $0.0546 |
| DeepSeek V4 Pro 0423 | Yes | 1,024,000 | $0.955 / $0.08 / $1.911 | $0.24 | $0.0672 |
| Gemini 3.8 Flash | No | 1,048,576 | $0.75 / $0.075 / $3.75 | $0.20 | $0.068 |
| GLM 5.3 | Yes | 1,048,576 | $1.40 / $0.26 / $4.40 | $0.36 | $0.14 |
| Kimi K3 | Yes | 1,048,576 | $0.53 / $0.30 / $12 | $0.19 | $0.15 |
| Claude Sonnet 5.5 | No | 1,000,000 | $2 / $0.10 / $10 | $0.53 | $0.18 |
| GPT-6.1 Sol | No | 1,050,000 | $2 / $0.10 / $10 | $0.53 | $0.18 |
| Gemini 3.1 Pro Preview | No | 1,048,576 | $2 / $0.20 / $12 | $0.54 | $0.19 |
| Grok 4.7 | No | 500,000 | $2 / $0.50 / $6 | $0.51 | $0.22 |
| Claude Opus 5.5 | No | 1,000,000 | $4 / $0.20 / $20 | $1.06 | $0.37 |
| Claude Fable 5.1 | No | 1,000,000 | $10 / $0.25 / $50 | $2.66 | $0.87 |
| GPT-6 Astra | No | 1,050,000 | $10 / $1 / $50 | $2.66 | $1.01 |
From our daily data. “Cached” assumes each step reads the previous prompt from the cache and pays the cache-write price, where one is listed, on new tokens. Open-weight prices are typical hosted prices and vary by host. A dash means no cache price is listed.
In this example, with caching applied wherever a model lists a cache price, the cheapest model costs $0.00975 per task (gpt-oss-120b) and the most expensive $1.01 (GPT-6 Astra). A cheaper model only saves money if it finishes the task in a similar number of steps: a model that needs twice as many steps, or two attempts, can cost more per finished task.
Free toolLLM cost calculatorIts “Coding agent” preset prices a large, mostly cached context per request across every model. Change the numbers to match your own agent’s usage.Tool-use reliability
An agent is only as good as its worst tool call. A malformed argument, the wrong tool, or a call that repeats forever wastes steps and money. Look for these, and test them on your own tools:
- Schema guarantees. Anthropic’s
strict: truemakes Claude’s tool calls always match your schema, and OpenAI’s strict mode makes function calls “reliably adhere to the function schema, instead of being best effort”. Turn them on where the provider supports them. - Tool count. Every definition is billed as input on every step. OpenAI suggests aiming for fewer than 20 functions at the start of a turn, and both providers offer tool search to load definitions on demand.
- Parallel calls. Reading three files at once saves steps. Both APIs support parallel tool calls and let you switch them off when order matters.
- Asking instead of guessing. Anthropic notes that Claude Opus is much more likely than Sonnet to ask for a missing required parameter rather than infer one, and that this is less dependable on less capable models.
The practical test: run 10 to 20 real tasks from your backlog and count malformed calls, wrong-tool calls and steps per task, not just whether the task passed.
Context window and output limit
Agents fill context quickly with file contents and tool output, so the window sets how long a task can run before the agent has to drop or summarise history. Of the 291 tool-capable text models in our data, 98 accept at least 1,000,000 tokens. Bigger isn’t free, though: each step pays for the whole context, and recall gets worse as context grows (see context windows explained).
The output limit matters when the agent writes whole files or long patches in one reply, and it includes any thinking. In our data, the limit per reply is 128,000 tokens on Claude Sonnet 5.5, 128,000 tokens on GPT-6.1 Sol and 65,536 tokens on Gemini 3.8 Flash. The context window checker shows whether a given file set fits with room for the reply.
For long tasks, compaction (summarising older steps on the server) and clearing old tool results keep the context small. Anthropic offers both for Claude, and OpenAI offers compaction in its Responses API.
Speed and reasoning modes
Speed decides how you use an agent: a fast model suits interactive pairing, a slow, thorough one suits tasks you hand off and review later. Anthropic labels its current lineup from “slower” (Fable) to “fastest” (Haiku), and both Anthropic and OpenAI sell faster processing for some models at a higher price.
218 of the 291 tool-capable models in our data can reason (think) before they act, and that thinking is billed as output. All three major providers let you set how much: effort on Claude, reasoning.effort on OpenAI and thinking_level on Gemini 3. A sensible pattern is a high setting for planning and debugging, and a low one, or a smaller model, for mechanical steps like renaming or formatting.
Where you can run it: APIs, clouds and open weights
If your company already buys cloud from one vendor, availability can decide the shortlist before quality does.
- Claude is sold through Anthropic’s API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Bedrock and Google Cloud set their own prices.
- OpenAI’s models are sold through OpenAI’s API, Microsoft Foundry on Azure and Amazon Bedrock.
- Gemini is sold through the Gemini API and Google Cloud’s Gemini Enterprise Agent Platform (formerly Vertex AI), which also hosts Claude and open models.
- Open-weight models can be downloaded and run on your own hardware or rented from many hosts. Our data lists 124 open-weight models with tool calling. Running a large one yourself needs a lot of GPU memory; the VRAM calculator estimates how much, and our guide to how much VRAM you need explains the maths.
What benchmarks like SWE-bench tell you
SWE-bench is a widely used coding-agent benchmark. Each task is a real GitHub issue: the system gets the issue text and a copy of the repository, must change the code to fix it, and passes only if tests that failed before the original fix now pass. The original set has 2,294 tasks from 12 popular Python repositories.
- SWE-bench Verified is a 500-task subset whose problem statements, tests and solvability were checked by human annotators, created with OpenAI.
- SWE-bench Multilingual has 300 tasks from 42 repositories in 9 programming languages, useful if you don’t write Python.
- Bash Only runs every model in the same minimal agent (mini-SWE-agent), so the scores compare models rather than agent harnesses. The site also charts results against cost and step count.
Use benchmarks to build a shortlist, not to make the final call. A leaderboard entry is a whole system (model plus agent), most tasks are Python, and your repository, tools and conventions are different. Scores also change as new models arrive, which is why we don’t quote them here.
A simple way to choose
- Shortlist three models in the model comparison: filter for tool calling and the context you need, then add one cheaper model and one stronger one.
- Run your own tasks. Take 10 to 20 closed issues from your backlog and run each through your agent on each model, with caching on.
- Measure per task: pass rate, cost (from the API’s usage figures), wall-clock time and failed tool calls.
- Pick the cheapest model that passes, and consider routing easy tasks to a smaller model.
- Recheck every few months. New models and price cuts arrive often; the prices on this page update daily.
FAQ
Questions people ask
What is the best model for coding agents?
There isn’t one that stays best for everyone. The right model is the cheapest one that reliably finishes your tasks with your tools and your codebase. Shortlist two or three from benchmark results and live prices, run 10 to 20 of your own tasks through each, and compare pass rate, cost per task and time.
Why are coding agents so expensive to run?
Because every step resends the whole conversation: system prompt, tool definitions, files read and test output. In our eight-step example the agent bills 237,300 input tokens to write 5,700. Prompt caching makes the re-sent part much cheaper, which is why it matters so much for agents.
Does prompt caching work with coding agents?
Yes, and it is the main way to cut their cost, because most of each request repeats the previous one. OpenAI and Gemini cache repeated prefixes automatically; on Claude you turn it on with cache_control. Watch the cache lifetime: Anthropic’s default is 5 minutes, so long pauses can make the next step pay full price.
Can I use an open-weight model for a coding agent?
Yes. Our data lists 124 open-weight models that support tool calling, available from hosted APIs or to run yourself. Self-hosting a large one needs a lot of GPU memory. Test tool calling on your own tasks with the exact host or serving setup you plan to use.
What does SWE-bench measure?
Whether a system can resolve real GitHub issues: it gets the issue and the repository, edits the code, and passes if tests that the original fix made pass now pass. Verified is a 500-task human-checked subset; Multilingual covers 9 languages. It measures model and agent together, on mostly Python projects.
Should I turn reasoning or thinking on for a coding agent?
For planning and debugging, usually yes; for mechanical edits, often no. Thinking tokens are billed as output and count toward the output limit, so a high setting on every step adds cost and time. Set the effort per task, or route simple steps to a cheaper model.
Try it
Tools from this guide
- AI Model ComparisonCompare prices, context windows and features across models.
- LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Context Window CheckerSee whether your text fits each model's context window.
- GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.
Keep reading
Related guides
- Tokens & costsContext windows explained: what counts, what happens at the limit, and how to check fit10 min read
- Tokens & costsPrompt caching explained: how it works on OpenAI, Anthropic and Gemini, and when it pays11 min read
- Coding agentsHow to reduce Claude Code token usage, and what each fix costs you11 min read