Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Guide · Coding agents

How to choose a model for coding agents: the criteria that matter

Choose a coding-agent model on the job it has to do, not on a leaderboard: reliable tool calls, enough context and output room, acceptable speed, and the lowest price per finished task. Agents resend their whole context on every step, so input tokens and prompt caching decide the bill far more than the output price does.

By Tahir NazirUpdated 10 min read

On this page
  1. What matters when choosing a model for a coding agent?
  2. Why agents cost more than chat, and how to price a task
  3. Tool-use reliability
  4. Context window and output limit
  5. Speed and reasoning modes
  6. Where you can run it: APIs, clouds and open weights
  7. What benchmarks like SWE-bench tell you
  8. A simple way to choose
  9. Questions people ask

What matters when choosing a model for a coding agent?

A coding agent is a loop: the model reads files, calls tools, edits code and runs tests until the task is done. The model you want is the cheapest one that finishes your tasks reliably inside that loop. Seven criteria decide it:

  1. Tool-use reliability: well-formed calls, the right tool, no loops.
  2. Context window: room for the repository slices, tool output and history a task accumulates.
  3. Output limit: room to write a large file or a long patch in one reply.
  4. Speed: how long a task takes end to end, which shapes how you work with the agent.
  5. Price per task: what a finished task costs, including caching, not the price per token.
  6. Reasoning controls: whether you can turn thinking up for hard steps and down for easy ones.
  7. Where it runs: the provider’s API, a cloud you already use, or open weights you host yourself.

Rankings change often, so this guide doesn’t name a winner. It shows how to judge each criterion and gives live prices for the current tool-capable models.

Why agents cost more than chat, and how to price a task

Every step of the loop is a new API request that resends everything so far: the system prompt, tool definitions, the task, every file read and every test log. Here is one small example task, a bug fix in eight steps:

Each request carries everything before it. Most of each bar was already sent in the previous step, which is exactly what prompt caching makes cheap.

The task bills 237,300 input tokens against 5,700 output tokens. Without caching, input is 89% of the cost on Claude Sonnet 5.5, even though its output price is 5× its input price. With caching, each step pays the cheap cached rate for everything it re-sent, and the bill for this task falls by 23% to 73% across the models below that list a cache price. A longer task has more steps, each resending more, so the gap grows.

Tool-capable models and what the example task costs (prices as of 2026-10-09), cheapest with caching first
ModelOpen weightsContext windowInput / cached / output per 1MTask, no cachingTask, cached
gpt-oss-120bYes131,072$0.037 / – / $0.17$0.00975$0.00975
Claude Haiku 5.5No1,000,000$0.10 / $0.01 / $0.50$0.0266$0.0101
GPT-6 LunaNo1,050,000$0.10 / $0.01 / $0.50$0.0266$0.0101
Qwen3.8 FlashYes1,000,000$0.15 / $0.016 / $0.47$0.0383$0.0143
DeepSeek V4.1 FlashYes1,048,576$0.30 / $0.006 / $1.20$0.078$0.0208
Mistral Large 4No1,048,576$0.68 / $0.07 / $2.09$0.17$0.0546
DeepSeek V4 Pro 0423Yes1,024,000$0.955 / $0.08 / $1.911$0.24$0.0672
Gemini 3.8 FlashNo1,048,576$0.75 / $0.075 / $3.75$0.20$0.068
GLM 5.3Yes1,048,576$1.40 / $0.26 / $4.40$0.36$0.14
Kimi K3Yes1,048,576$0.53 / $0.30 / $12$0.19$0.15
Claude Sonnet 5.5No1,000,000$2 / $0.10 / $10$0.53$0.18
GPT-6.1 SolNo1,050,000$2 / $0.10 / $10$0.53$0.18
Gemini 3.1 Pro PreviewNo1,048,576$2 / $0.20 / $12$0.54$0.19
Grok 4.7No500,000$2 / $0.50 / $6$0.51$0.22
Claude Opus 5.5No1,000,000$4 / $0.20 / $20$1.06$0.37
Claude Fable 5.1No1,000,000$10 / $0.25 / $50$2.66$0.87
GPT-6 AstraNo1,050,000$10 / $1 / $50$2.66$1.01

From our daily data. “Cached” assumes each step reads the previous prompt from the cache and pays the cache-write price, where one is listed, on new tokens. Open-weight prices are typical hosted prices and vary by host. A dash means no cache price is listed.

In this example, with caching applied wherever a model lists a cache price, the cheapest model costs $0.00975 per task (gpt-oss-120b) and the most expensive $1.01 (GPT-6 Astra). A cheaper model only saves money if it finishes the task in a similar number of steps: a model that needs twice as many steps, or two attempts, can cost more per finished task.

Free toolLLM cost calculatorIts “Coding agent” preset prices a large, mostly cached context per request across every model. Change the numbers to match your own agent’s usage.

Tool-use reliability

An agent is only as good as its worst tool call. A malformed argument, the wrong tool, or a call that repeats forever wastes steps and money. Look for these, and test them on your own tools:

  • Schema guarantees. Anthropic’s strict: true makes Claude’s tool calls always match your schema, and OpenAI’s strict mode makes function calls “reliably adhere to the function schema, instead of being best effort”. Turn them on where the provider supports them.
  • Tool count. Every definition is billed as input on every step. OpenAI suggests aiming for fewer than 20 functions at the start of a turn, and both providers offer tool search to load definitions on demand.
  • Parallel calls. Reading three files at once saves steps. Both APIs support parallel tool calls and let you switch them off when order matters.
  • Asking instead of guessing. Anthropic notes that Claude Opus is much more likely than Sonnet to ask for a missing required parameter rather than infer one, and that this is less dependable on less capable models.

The practical test: run 10 to 20 real tasks from your backlog and count malformed calls, wrong-tool calls and steps per task, not just whether the task passed.

Context window and output limit

Agents fill context quickly with file contents and tool output, so the window sets how long a task can run before the agent has to drop or summarise history. Of the 291 tool-capable text models in our data, 98 accept at least 1,000,000 tokens. Bigger isn’t free, though: each step pays for the whole context, and recall gets worse as context grows (see context windows explained).

The output limit matters when the agent writes whole files or long patches in one reply, and it includes any thinking. In our data, the limit per reply is 128,000 tokens on Claude Sonnet 5.5, 128,000 tokens on GPT-6.1 Sol and 65,536 tokens on Gemini 3.8 Flash. The context window checker shows whether a given file set fits with room for the reply.

For long tasks, compaction (summarising older steps on the server) and clearing old tool results keep the context small. Anthropic offers both for Claude, and OpenAI offers compaction in its Responses API.

Speed and reasoning modes

Speed decides how you use an agent: a fast model suits interactive pairing, a slow, thorough one suits tasks you hand off and review later. Anthropic labels its current lineup from “slower” (Fable) to “fastest” (Haiku), and both Anthropic and OpenAI sell faster processing for some models at a higher price.

218 of the 291 tool-capable models in our data can reason (think) before they act, and that thinking is billed as output. All three major providers let you set how much: effort on Claude, reasoning.effort on OpenAI and thinking_level on Gemini 3. A sensible pattern is a high setting for planning and debugging, and a low one, or a smaller model, for mechanical steps like renaming or formatting.

Where you can run it: APIs, clouds and open weights

If your company already buys cloud from one vendor, availability can decide the shortlist before quality does.

  • Claude is sold through Anthropic’s API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Bedrock and Google Cloud set their own prices.
  • OpenAI’s models are sold through OpenAI’s API, Microsoft Foundry on Azure and Amazon Bedrock.
  • Gemini is sold through the Gemini API and Google Cloud’s Gemini Enterprise Agent Platform (formerly Vertex AI), which also hosts Claude and open models.
  • Open-weight models can be downloaded and run on your own hardware or rented from many hosts. Our data lists 124 open-weight models with tool calling. Running a large one yourself needs a lot of GPU memory; the VRAM calculator estimates how much, and our guide to how much VRAM you need explains the maths.

What benchmarks like SWE-bench tell you

SWE-bench is a widely used coding-agent benchmark. Each task is a real GitHub issue: the system gets the issue text and a copy of the repository, must change the code to fix it, and passes only if tests that failed before the original fix now pass. The original set has 2,294 tasks from 12 popular Python repositories.

  • SWE-bench Verified is a 500-task subset whose problem statements, tests and solvability were checked by human annotators, created with OpenAI.
  • SWE-bench Multilingual has 300 tasks from 42 repositories in 9 programming languages, useful if you don’t write Python.
  • Bash Only runs every model in the same minimal agent (mini-SWE-agent), so the scores compare models rather than agent harnesses. The site also charts results against cost and step count.

Use benchmarks to build a shortlist, not to make the final call. A leaderboard entry is a whole system (model plus agent), most tasks are Python, and your repository, tools and conventions are different. Scores also change as new models arrive, which is why we don’t quote them here.

A simple way to choose

  1. Shortlist three models in the model comparison: filter for tool calling and the context you need, then add one cheaper model and one stronger one.
  2. Run your own tasks. Take 10 to 20 closed issues from your backlog and run each through your agent on each model, with caching on.
  3. Measure per task: pass rate, cost (from the API’s usage figures), wall-clock time and failed tool calls.
  4. Pick the cheapest model that passes, and consider routing easy tasks to a smaller model.
  5. Recheck every few months. New models and price cuts arrive often; the prices on this page update daily.

FAQ

Questions people ask

What is the best model for coding agents?

There isn’t one that stays best for everyone. The right model is the cheapest one that reliably finishes your tasks with your tools and your codebase. Shortlist two or three from benchmark results and live prices, run 10 to 20 of your own tasks through each, and compare pass rate, cost per task and time.

Why are coding agents so expensive to run?

Because every step resends the whole conversation: system prompt, tool definitions, files read and test output. In our eight-step example the agent bills 237,300 input tokens to write 5,700. Prompt caching makes the re-sent part much cheaper, which is why it matters so much for agents.

Does prompt caching work with coding agents?

Yes, and it is the main way to cut their cost, because most of each request repeats the previous one. OpenAI and Gemini cache repeated prefixes automatically; on Claude you turn it on with cache_control. Watch the cache lifetime: Anthropic’s default is 5 minutes, so long pauses can make the next step pay full price.

Can I use an open-weight model for a coding agent?

Yes. Our data lists 124 open-weight models that support tool calling, available from hosted APIs or to run yourself. Self-hosting a large one needs a lot of GPU memory. Test tool calling on your own tasks with the exact host or serving setup you plan to use.

What does SWE-bench measure?

Whether a system can resolve real GitHub issues: it gets the issue and the repository, edits the code, and passes if tests that the original fix made pass now pass. Verified is a 500-task human-checked subset; Multilingual covers 9 languages. It measures model and agent together, on mostly Python projects.

Should I turn reasoning or thinking on for a coding agent?

For planning and debugging, usually yes; for mechanical edits, often no. Thinking tokens are billed as output and count toward the output limit, so a high setting on every step adds cost and time. Set the effort per task, or route simple steps to a cheaper model.

Try it

Tools from this guide

Keep reading