Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Guide · Tokens & costs

The cheapest LLM APIs right now, ranked from daily price data

By blended price (three input tokens for every output token), the cheapest LLM API in our data on 2026-10-09 is Mistral Nemo from Mistral, at $0.0293 per million tokens. The ranking is rebuilt from public price lists whenever the daily refresh finds a change, and the cheapest model for your app depends on its mix of input, output, caching and quality needs.

By Tahir NazirUpdated 9 min read

On this page
  1. What is the cheapest LLM API right now?
  2. How to compare LLM prices: blended price explained
  3. The cheapest LLM for tools, vision, reasoning and long context
  4. Why the cheapest per token isn’t always the cheapest per task
  5. Is there a free LLM API?
  6. Batch and caching discounts
  7. Open-weight models through hosted APIs
  8. How to choose the cheapest model that works
  9. Questions people ask

What is the cheapest LLM API right now?

On 2026-10-09, the cheapest model in our data by blended price is Mistral Nemo (Mistral): $0.029 per million input tokens and $0.03 per million output tokens. These are the ten cheapest of the 317 text models we rank:

Blended price per million tokens, weighting input three to one. The median of all 317 ranked models is $0.76. Rebuilt from the daily data on every deploy.

For scale, Anthropic’s mid-tier Claude Sonnet 5.5 has a blended price 137 times that of Mistral Nemo, and the most expensive model we rank, at $262.50, is about 9,000 times. Price alone says nothing about quality, so treat the ladder as a shortlist, not a recommendation.

How to compare LLM prices: blended price explained

Every API has two prices, input and output, so comparing models needs one number that combines them. A blended price is a weighted average for an assumed mix of tokens. We use three input tokens for every output token, a common shape for chat, RAG and classification, where prompts are longer than answers:

Blended price (USD per 1M tokens)
blended = (3 × input price + 1 × output price) ÷ 4

The weighting matters. Long-form writing, code generation and reasoning models produce far more output, which favours models with cheap output. Here is the cheapest model under four different weightings:

Cheapest model by input:output weighting
WeightingCheapest modelBlended price
Input onlyGranite 4.0 Micro$0.017 per 1M
3 input : 1 output (our ranking)Mistral Nemo$0.0293 per 1M
1 : 1Mistral Nemo$0.0295 per 1M
1 input : 3 outputMistral Nemo$0.0297 per 1M

Two different models win across the four weightings today, and eight of the ten cheapest at 3:1 are still in the ten cheapest at 1:3. Prices as of 2026-10-09.

Blended prices use each model’s base rate. Some models charge more once a prompt passes a size threshold, and a few list introductory prices with an end date, so check the provider’s page before you commit a budget.

The cheapest LLM for tools, vision, reasoning and long context

Most apps have hard requirements before price comes into it. The cheapest model that meets each one:

Cheapest model for common requirements
RequirementCheapest modelInput / output per 1MBlended
Any text modelMistral Nemo$0.029 / $0.03$0.0293
Tool (function) callingMistral Nemo$0.029 / $0.03$0.0293
Image inputLing 3.0 Flash VL$0.021 / $0.0616$0.0312
ReasoningLing 3.0 Flash VL$0.021 / $0.0616$0.0312
1M-token context windowQwen3.7 Flash$0.03 / $0.13$0.055
Open weightsMistral Nemo$0.029 / $0.03$0.0293
Not open weightsQwen3.7 Flash$0.03 / $0.13$0.055
At batch pricesgpt-oss-20b$0.018 / $0.09$0.046 (batch)
From OpenAIGPT-5 Nano$0.05 / $0.40$0.1375
From AnthropicClaude Haiku 5.5$0.10 / $0.50$0.20
From GoogleGemini 2.5 Flash Lite$0.10 / $0.40$0.175

Capabilities as reported by the data sources; check the provider’s docs for details such as how many images or tools a model accepts. Prices as of 2026-10-09.

The model comparison filters every model by context window, tool calling, image input, reasoning and prompt caching, and sorts by price, so you can go further down each list than the single winner shown here.

Why the cheapest per token isn’t always the cheapest per task

Your bill is price × tokens, and the token side varies between models as much as the price side. Four things decide the real cost of a task:

  • Tokenizers. Each model family splits text differently, so the same prompt is a different number of tokens. Anthropic says the tokenizer in Claude 4.7 and later produces about 30% more tokens for the same text than its previous one. The token counter shows the count per model, and what is a token explains why.
  • Verbosity. A model at half the price per token that writes twice as much costs the same per answer. Measure output length on your own prompts.
  • Reasoning tokens. Reasoning models bill their hidden thinking as output tokens, which can be many times the visible answer.
  • Failures and retries. A cheap model that returns invalid JSON or a wrong answer one time in five needs a retry, a repair step or a fallback to a bigger model. Cost per successful task is the number that matters.

Workload shape changes the ranking too. Running the cost calculator’s four presets (with their caching and batch settings) across every model that fits gives these winners:

Cheapest model for four workloads
WorkloadCheapest modelPer month
Support chatbotLing 3.0 Flash VL$1.32
RAG appLing 3.0 Flash VL$2.08
Coding agentLing 3.0 Flash VL$2.59
Document summariserGranite 4.0 Micro$1.85

Each preset’s token counts, requests per day, cached share and batch setting are shown in the [cost calculator](/llm-cost-calculator). Prices as of 2026-10-09.

Today that gives two different winners for four workloads. Cached-input prices, batch prices and context limits all move the ranking.

Free toolLLM API cost calculatorEnter your own tokens per request and volume to see the 10 cheapest models that fit your workload, with caching and batch pricing applied.

Is there a free LLM API?

Yes, with limits. These are the free options we verified in the providers’ own documentation on 8 October 2026:

  • Google Gemini free tier. On 2026-10-08, Google’s pricing page listed text models including Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite and Gemini 3.1 Flash-Lite as free of charge for input and output. Rate limits apply per project, and Google says free-tier content is used to improve its products, so keep personal and confidential data off it. How to get a free Gemini API key walks through it.
  • OpenRouter free variants. OpenRouter serves some models as free variants (IDs ending in :free), limited to 20 requests a minute and 50 requests a day, or 1,000 a day once you have bought at least 10 credits.
  • Anthropic trial credits. Anthropic says new users receive a small amount of free credits to test the API. That’s a trial, not a free tier.

Free tiers suit prototypes, demos and personal tools. For production, plan on paid prices: limits are low and terms can change. You can try the Gemini models free in our Gemini playground with your own key.

Batch and caching discounts

A model’s list price isn’t its lowest price. Two discounts often matter more than the gap between two budget models:

  • Batch APIs. OpenAI, Anthropic and Google charge 50% less for requests processed asynchronously (OpenAI promises results within 24 hours). Batch prices are listed for 82 of the 317 models we rank, and backfills, nightly jobs and evaluations rarely need an instant answer.
  • Prompt caching. Repeated prefixes (system prompts, tool definitions, documents) are billed at a fraction of the input price on models that support caching. For agents and document Q&A, the cached-input price can matter more than the list input price. See prompt caching explained.

That’s why the four workloads above don’t all share a winner with the ladder, and why the cost estimation guide prices whole workloads rather than single requests.

Open-weight models through hosted APIs

Open-weight models (Llama, Qwen, Mistral, DeepSeek, gpt-oss and others) publish their weights, so many companies can host them and compete on price. That competition helps explain why nine of the ten cheapest models today are open-weight. The price we show for them is a typical hosted price from our data sources, not necessarily the lowest any host charges.

  • Prices differ by host. OpenRouter routes each request among the hosts that serve a model and, by default, prefers cheaper ones (it weights the choice by the inverse square of the price). Setting provider.sort to "price", or adding :floor to the model ID, asks for the cheapest host.
  • Hosts differ in more than price. OpenRouter notes that providers serve models at different quantisation levels (such as INT4, INT8 or FP8) to cut compute, so the same model can give different results from one host to another.
  • Self-hosting is a different calculation. Running a model on your own GPUs swaps per-token prices for hardware and electricity. The VRAM calculator shows what hardware a model needs.

How to choose the cheapest model that works

  1. Write down the hard requirements: context size, tool calling, image input, structured output, data location.
  2. Shortlist three to five of the cheapest models that meet them, using the table above or the model comparison.
  3. Run 50 to 100 real prompts through each, and record quality, input and output tokens, reasoning tokens and failures.
  4. Compare cost per successful task, including retries and any fallback to a larger model.
  5. Apply batch and caching where your workload allows, then re-check every month or two. Prices in this market change often; this page is rebuilt from the daily data on every deploy.

FAQ

Questions people ask

What is the cheapest LLM API?

By blended price (three input tokens per output token), the cheapest text model in our data on 2026-10-09 is Mistral Nemo from Mistral, at $0.029 per million input tokens and $0.03 per million output tokens. The cheapest model for your app depends on your token mix, required features and how often a cheap model needs a retry.

What is the cheapest OpenAI model?

Among OpenAI’s own API models in our data, the cheapest by blended price on 2026-10-09 is GPT-5 Nano, at $0.05 input and $0.40 output per million tokens. OpenAI’s Batch API halves those prices for work that can wait up to 24 hours.

What is the cheapest Claude model?

The cheapest Claude model in our data on 2026-10-09 is Claude Haiku 5.5, at $0.10 input and $0.50 output per million tokens. Anthropic’s Message Batches API halves that price, and prompt caching cuts the cost of repeated input.

What is the cheapest Gemini model?

By list price, the cheapest Gemini model in our data on 2026-10-09 is Gemini 2.5 Flash Lite, at $0.10 input and $0.40 output per million tokens. Several Gemini models are also free on Google’s free tier, within rate limits and with your content used to improve Google’s products.

Are open-source models cheaper through an API?

Often, yes: nine of the ten cheapest models we rank today are open-weight, because many hosting companies compete to serve them. The same model can cost different amounts on different hosts, and hosts may run different quantisations, so compare quality as well as price.

How often do LLM API prices change?

Often enough that any fixed list goes stale within weeks. Our prices are refreshed every day from the OpenRouter models API and the LiteLLM price list (last refreshed 2026-10-09), and this ranking is rebuilt from that data. Always confirm the price on the provider’s pricing page before committing a budget.

Try it

Tools from this guide

Keep reading