Guide · Tokens & costs
The cheapest LLM APIs right now, ranked from daily price data
By blended price (three input tokens for every output token), the cheapest LLM API in our data on 2026-10-09 is Mistral Nemo from Mistral, at $0.0293 per million tokens. The ranking is rebuilt from public price lists whenever the daily refresh finds a change, and the cheapest model for your app depends on its mix of input, output, caching and quality needs.
By Tahir NazirUpdated 9 min read
On this page
- What is the cheapest LLM API right now?
- How to compare LLM prices: blended price explained
- The cheapest LLM for tools, vision, reasoning and long context
- Why the cheapest per token isn’t always the cheapest per task
- Is there a free LLM API?
- Batch and caching discounts
- Open-weight models through hosted APIs
- How to choose the cheapest model that works
- Questions people ask
What is the cheapest LLM API right now?
On 2026-10-09, the cheapest model in our data by blended price is Mistral Nemo (Mistral): $0.029 per million input tokens and $0.03 per million output tokens. These are the ten cheapest of the 317 text models we rank:
- 1Mistral NemoMistral · $0.029 in · $0.03 out$0.0293
- 2Ling 3.0 Flash VLinclusionAI · $0.021 in · $0.0616 out$0.0312
- 3Ling 3.0 FlashinclusionAI · $0.021 in · $0.063 out$0.0315
- 4gpt-oss-20bOpenAI · $0.018 in · $0.09 out$0.036
- 5Granite 4.0 MicroIBM · $0.017 in · $0.112 out$0.0408
- 6Nex-N2.5-MiniNex AGI · $0.025 in · $0.10 out$0.0438
- 7Qwen3.7 FlashQwen · $0.03 in · $0.13 out$0.055
- 8Llama 3.1 8B InstructMeta · $0.05 in · $0.08 out$0.0575
- 9Mistral Small 3Mistral · $0.05 in · $0.08 out$0.0575
- 10Schematron V2 TurboInference.net · $0.03 in · $0.15 out$0.06
Blended = (3 × input + 1 × output) ÷ 4, USD per 1M tokens. Median of all 317 ranked models: $0.76.
Prices as of 2026-10-09
For scale, Anthropic’s mid-tier Claude Sonnet 5.5 has a blended price 137 times that of Mistral Nemo, and the most expensive model we rank, at $262.50, is about 9,000 times. Price alone says nothing about quality, so treat the ladder as a shortlist, not a recommendation.
How to compare LLM prices: blended price explained
Every API has two prices, input and output, so comparing models needs one number that combines them. A blended price is a weighted average for an assumed mix of tokens. We use three input tokens for every output token, a common shape for chat, RAG and classification, where prompts are longer than answers:
blended = (3 × input price + 1 × output price) ÷ 4The weighting matters. Long-form writing, code generation and reasoning models produce far more output, which favours models with cheap output. Here is the cheapest model under four different weightings:
| Weighting | Cheapest model | Blended price |
|---|---|---|
| Input only | Granite 4.0 Micro | $0.017 per 1M |
| 3 input : 1 output (our ranking) | Mistral Nemo | $0.0293 per 1M |
| 1 : 1 | Mistral Nemo | $0.0295 per 1M |
| 1 input : 3 output | Mistral Nemo | $0.0297 per 1M |
Two different models win across the four weightings today, and eight of the ten cheapest at 3:1 are still in the ten cheapest at 1:3. Prices as of 2026-10-09.
Blended prices use each model’s base rate. Some models charge more once a prompt passes a size threshold, and a few list introductory prices with an end date, so check the provider’s page before you commit a budget.
The cheapest LLM for tools, vision, reasoning and long context
Most apps have hard requirements before price comes into it. The cheapest model that meets each one:
| Requirement | Cheapest model | Input / output per 1M | Blended |
|---|---|---|---|
| Any text model | Mistral Nemo | $0.029 / $0.03 | $0.0293 |
| Tool (function) calling | Mistral Nemo | $0.029 / $0.03 | $0.0293 |
| Image input | Ling 3.0 Flash VL | $0.021 / $0.0616 | $0.0312 |
| Reasoning | Ling 3.0 Flash VL | $0.021 / $0.0616 | $0.0312 |
| 1M-token context window | Qwen3.7 Flash | $0.03 / $0.13 | $0.055 |
| Open weights | Mistral Nemo | $0.029 / $0.03 | $0.0293 |
| Not open weights | Qwen3.7 Flash | $0.03 / $0.13 | $0.055 |
| At batch prices | gpt-oss-20b | $0.018 / $0.09 | $0.046 (batch) |
| From OpenAI | GPT-5 Nano | $0.05 / $0.40 | $0.1375 |
| From Anthropic | Claude Haiku 5.5 | $0.10 / $0.50 | $0.20 |
| From Google | Gemini 2.5 Flash Lite | $0.10 / $0.40 | $0.175 |
Capabilities as reported by the data sources; check the provider’s docs for details such as how many images or tools a model accepts. Prices as of 2026-10-09.
The model comparison filters every model by context window, tool calling, image input, reasoning and prompt caching, and sorts by price, so you can go further down each list than the single winner shown here.
Why the cheapest per token isn’t always the cheapest per task
Your bill is price × tokens, and the token side varies between models as much as the price side. Four things decide the real cost of a task:
- Tokenizers. Each model family splits text differently, so the same prompt is a different number of tokens. Anthropic says the tokenizer in Claude 4.7 and later produces about 30% more tokens for the same text than its previous one. The token counter shows the count per model, and what is a token explains why.
- Verbosity. A model at half the price per token that writes twice as much costs the same per answer. Measure output length on your own prompts.
- Reasoning tokens. Reasoning models bill their hidden thinking as output tokens, which can be many times the visible answer.
- Failures and retries. A cheap model that returns invalid JSON or a wrong answer one time in five needs a retry, a repair step or a fallback to a bigger model. Cost per successful task is the number that matters.
Workload shape changes the ranking too. Running the cost calculator’s four presets (with their caching and batch settings) across every model that fits gives these winners:
| Workload | Cheapest model | Per month |
|---|---|---|
| Support chatbot | Ling 3.0 Flash VL | $1.32 |
| RAG app | Ling 3.0 Flash VL | $2.08 |
| Coding agent | Ling 3.0 Flash VL | $2.59 |
| Document summariser | Granite 4.0 Micro | $1.85 |
Each preset’s token counts, requests per day, cached share and batch setting are shown in the [cost calculator](/llm-cost-calculator). Prices as of 2026-10-09.
Today that gives two different winners for four workloads. Cached-input prices, batch prices and context limits all move the ranking.
Free toolLLM API cost calculatorEnter your own tokens per request and volume to see the 10 cheapest models that fit your workload, with caching and batch pricing applied.Is there a free LLM API?
Yes, with limits. These are the free options we verified in the providers’ own documentation on 8 October 2026:
- Google Gemini free tier. On 2026-10-08, Google’s pricing page listed text models including Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite and Gemini 3.1 Flash-Lite as free of charge for input and output. Rate limits apply per project, and Google says free-tier content is used to improve its products, so keep personal and confidential data off it. How to get a free Gemini API key walks through it.
- OpenRouter free variants. OpenRouter serves some models as free variants (IDs ending in
:free), limited to 20 requests a minute and 50 requests a day, or 1,000 a day once you have bought at least 10 credits. - Anthropic trial credits. Anthropic says new users receive a small amount of free credits to test the API. That’s a trial, not a free tier.
Free tiers suit prototypes, demos and personal tools. For production, plan on paid prices: limits are low and terms can change. You can try the Gemini models free in our Gemini playground with your own key.
Batch and caching discounts
A model’s list price isn’t its lowest price. Two discounts often matter more than the gap between two budget models:
- Batch APIs. OpenAI, Anthropic and Google charge 50% less for requests processed asynchronously (OpenAI promises results within 24 hours). Batch prices are listed for 82 of the 317 models we rank, and backfills, nightly jobs and evaluations rarely need an instant answer.
- Prompt caching. Repeated prefixes (system prompts, tool definitions, documents) are billed at a fraction of the input price on models that support caching. For agents and document Q&A, the cached-input price can matter more than the list input price. See prompt caching explained.
That’s why the four workloads above don’t all share a winner with the ladder, and why the cost estimation guide prices whole workloads rather than single requests.
Open-weight models through hosted APIs
Open-weight models (Llama, Qwen, Mistral, DeepSeek, gpt-oss and others) publish their weights, so many companies can host them and compete on price. That competition helps explain why nine of the ten cheapest models today are open-weight. The price we show for them is a typical hosted price from our data sources, not necessarily the lowest any host charges.
- Prices differ by host. OpenRouter routes each request among the hosts that serve a model and, by default, prefers cheaper ones (it weights the choice by the inverse square of the price). Setting
provider.sortto"price", or adding:floorto the model ID, asks for the cheapest host. - Hosts differ in more than price. OpenRouter notes that providers serve models at different quantisation levels (such as INT4, INT8 or FP8) to cut compute, so the same model can give different results from one host to another.
- Self-hosting is a different calculation. Running a model on your own GPUs swaps per-token prices for hardware and electricity. The VRAM calculator shows what hardware a model needs.
How to choose the cheapest model that works
- Write down the hard requirements: context size, tool calling, image input, structured output, data location.
- Shortlist three to five of the cheapest models that meet them, using the table above or the model comparison.
- Run 50 to 100 real prompts through each, and record quality, input and output tokens, reasoning tokens and failures.
- Compare cost per successful task, including retries and any fallback to a larger model.
- Apply batch and caching where your workload allows, then re-check every month or two. Prices in this market change often; this page is rebuilt from the daily data on every deploy.
FAQ
Questions people ask
What is the cheapest LLM API?
By blended price (three input tokens per output token), the cheapest text model in our data on 2026-10-09 is Mistral Nemo from Mistral, at $0.029 per million input tokens and $0.03 per million output tokens. The cheapest model for your app depends on your token mix, required features and how often a cheap model needs a retry.
What is the cheapest OpenAI model?
Among OpenAI’s own API models in our data, the cheapest by blended price on 2026-10-09 is GPT-5 Nano, at $0.05 input and $0.40 output per million tokens. OpenAI’s Batch API halves those prices for work that can wait up to 24 hours.
What is the cheapest Claude model?
The cheapest Claude model in our data on 2026-10-09 is Claude Haiku 5.5, at $0.10 input and $0.50 output per million tokens. Anthropic’s Message Batches API halves that price, and prompt caching cuts the cost of repeated input.
What is the cheapest Gemini model?
By list price, the cheapest Gemini model in our data on 2026-10-09 is Gemini 2.5 Flash Lite, at $0.10 input and $0.40 output per million tokens. Several Gemini models are also free on Google’s free tier, within rate limits and with your content used to improve Google’s products.
Are open-source models cheaper through an API?
Often, yes: nine of the ten cheapest models we rank today are open-weight, because many hosting companies compete to serve them. The same model can cost different amounts on different hosts, and hosts may run different quantisations, so compare quality as well as price.
How often do LLM API prices change?
Often enough that any fixed list goes stale within weeks. Our prices are refreshed every day from the OpenRouter models API and the LiteLLM price list (last refreshed 2026-10-09), and this ranking is rebuilt from that data. Always confirm the price on the provider’s pricing page before committing a budget.
Try it
Tools from this guide
Keep reading