Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Tokens & Costs

AI model comparison: prices, context windows and features of 340+ LLMs

Compare every major AI model's API price, context window, output limit and capabilities in one sortable table, updated daily.

Providers
Must support

Showing 50 of 333 matching models (333 in total)

Prices updated 2026-10-09
AI model prices in US dollars per million tokens, context windows and release dates. Column headers sort the table.
Cost estimate
Step 5 PreviewStepFun$1$2.70$0.051M64K$1.002026-10-08Estimate cost
Claude Haiku 5.5Anthropic$0.10$0.50$0.011M128K$0.502026-10-07Estimate cost
Mistral Large 4Mistral$0.68$2.09$0.071.05M262K$0.712026-10-06Estimate cost
Ling 3.1 FlashinclusionAI$0$0–262K33K$0.002026-10-02Estimate cost
Pareto 26.10 Previewunbiased$0.80$3.20$0.031.05M131K$0.842026-10-01Estimate cost
GPT-6.1 SolOpenAI$2$10$0.101.05M128K$4.202026-09-29Estimate cost
GPT-6.1 Sol ProOpenAI$2$10$0.101.05M128K$4.202026-09-29Estimate cost
Claude Sonnet 5.5Anthropic$2$10$0.101M128K$2.002026-09-28Estimate cost
Perceptron Mk1.5Perceptron$0.15$1.50–37K8K$0.005532026-09-25Estimate cost
Ember-1Fireworks$3$15$0.301.05M–$3.152026-09-24Estimate cost
Aion 3.5AionLabs$3$6$0.75262K33K$0.792026-09-23Estimate cost
Aion 3.5 MiniAionLabs$0.70$1.40$0.18262K33K$0.182026-09-23Estimate cost
GLM 5.3 PrimeZ.ai$2.80$8.80$0.561M131K$2.802026-09-23Estimate cost
Qwen3.8 Max PrimeQwen$4$12$0.501M131K$4.002026-09-23Estimate cost
Solar Mini 4Upstage$0.05$0.20$0.005524K131K$0.02622026-09-23Estimate cost
Claude Opus 5.5Anthropic$4$20$0.201M128K$4.002026-09-22Estimate cost
Command A+Cohere$0.30$1.50$0.15192K64K$0.05762026-09-22Estimate cost
GPT-6 LunaOpenAI$0.10$0.50$0.011.05M128K$0.212026-09-22Estimate cost
GPT-6 Luna ProOpenAI$0.10$0.50$0.011.05M128K$0.212026-09-22Estimate cost
GPT-6 SolOpenAI$2$10$0.201.05M128K$4.202026-09-22Estimate cost
GPT-6 Sol ProOpenAI$2$10$0.201.05M128K$4.202026-09-22Estimate cost
Grok 4.7xAI$2$6$0.50500K–$2.002026-09-21Estimate cost
MiMo-V2.6-FlashXiaomiOpen$0.14$0.28$0.00281.05M131K$0.152026-09-21Estimate cost
MiMo-V2.6-ProXiaomiOpen$0.435$0.87$0.00361.05M131K$0.462026-09-21Estimate cost
MiMo-V2.6-Pro-UltraSpeedXiaomi$4.35$8.70$0.0361.05M131K$4.562026-09-21Estimate cost
Qwen3.8 Omni FlashQwen$0.15$0.47$0.0161M131K$0.152026-09-21Estimate cost
GLM 5.3 FlashXZ.ai$0.37$1.25$0.091.05M131K$0.392026-09-18Estimate cost
Ternary Bonsai 2 27BPrismMLOpen$0.075$0.50$0.0375262K33K$0.01972026-09-18Estimate cost
Paretounbiased$2.50$7.50$0.25262K131K$0.662026-09-17Estimate cost
Schematron V2 SmallInference.netOpen$0.05$0.23$0.05128K4K$0.00642026-09-12Estimate cost
Schematron V2 TurboInference.netOpen$0.03$0.15$0.03128K8K$0.003842026-09-12Estimate cost
Fugu MaxSakana$2$6$0.251M128K$2.002026-09-11Estimate cost
Fugu Ultra v2Sakana$5$30$0.501M128K$10.002026-09-11Estimate cost
DeepSeek V4.1 FlashDeepSeekOpen$0.30$1.20$0.0061.05M–$0.312026-09-10Estimate cost
Ling 3.0 Flash VLinclusionAIOpen$0.021$0.0616$0.0042262K33K$0.005512026-09-10Estimate cost
Mercury 2.5Inception$0.04$0.15$0.004260K66K$0.01042026-09-08Estimate cost
Nex-N2.5-MiniNex AGIOpen$0.025$0.10$0.0025262K–$0.006552026-09-08Estimate cost
Nex-N2.5-ProNex AGIOpen$0.075$0.25$0.015262K–$0.01972026-09-08Estimate cost
GPT-6 AstraOpenAI$10$50$11.05M128K$21.002026-09-04Estimate cost
GPT-6 Astra ProOpenAI$10$50$11.05M128K$21.002026-09-04Estimate cost
Ling 3.0 Flash SanteinclusionAI$0.042$0.1232$0.0084262K33K$0.0112026-09-04Estimate cost
Qwen3.8 Max (0902)Qwen$2$6$0.251M131K$2.002026-09-03Estimate cost
Gemini 3.8 FlashGoogle$0.75$3.75$0.0751.05M66K$0.792026-09-02Estimate cost
Muse Spark 1.3Meta$1.25$4.25$0.151.05M–$1.312026-09-02Estimate cost
Muse Spark 1.3 ContributorMeta$0.10$0.20$0.0021.05M–$0.102026-09-02Estimate cost
Claude Fable 5.1Anthropic$10$50$0.251M128K$10.002026-09-01Estimate cost
Granite 4.2 8BIBMOpen$0.06$0.25$0.015131K–$0.007862026-08-31Estimate cost
Ling 3.0 Flash FininclusionAI$0.042$0.1232$0.0084262K33K$0.0112026-08-27Estimate cost
GLM 5.3 FlashZ.aiOpen$0.15$0.50$0.031.05M–$0.162026-08-26Estimate cost
Qwen3.8 FlashQwenOpen$0.15$0.47$0.0161M131K$0.152026-08-26Estimate cost

Steps

How to use the AI model comparison

  1. Search for a model, or pick one or more providers.
  2. Set a minimum context window and any capabilities you need, such as image input or tool calling.
  3. Click a column header to sort, for example by input price or context window. Click again to reverse.
  4. Use “Estimate cost” on any row to price your own workload in the LLM cost calculator.
  5. Copy the page address to share the exact filters and sort you’re looking at.

Method

How it works

Choosing a model usually starts with three questions: can it handle my input, can it do what I need, and what will it cost? This table answers all three for 340+ models at once, so you can narrow the field before testing anything.

Reading the prices

Prices are list prices in US dollars per million tokens. Input covers everything you send (system prompt, history, documents); output is the model’s reply and costs more because it’s generated one token at a time. Cached input is the discounted price for re-reading a prompt prefix the provider has already cached, which matters for chatbots and agents that resend the same instructions on every request. A dash means the provider hasn’t published that price.

Context window, max output and the cost of filling it

The context window is the most a single request can hold, prompt and reply together. Max output limits the reply on its own. The Fill context column multiplies the window by the input price, which tells you what one maximum-length prompt costs. Two models can share a 1M-token window and still differ by more than a hundred times on this column, so it’s a quick way to see which long-context models are practical for everyday use.

Capabilities and open weights

The capability filters (image input, tool calling, reasoning, prompt caching, structured output) use the features each provider declares for its API. They tell you a feature exists, not how well it works. Open-weight models publish their weights, so several companies host them and you can also run them yourself; their price here is a typical hosted price, and our VRAM calculator covers self-hosting.

Where the data comes from

The table is generated from the OpenRouter models API and the open-source LiteLLM price list, refreshed every day. We correct the data by hand when a source is wrong; for example, when a feed reports a cache price as the normal input price, the model is held back until it’s checked against the provider’s pricing page. Release dates are when a model was first listed by our source, which can differ from the announcement date. The methodology page has the details.

Examples

Worked examples

Largest context windows

Most tokens a single request can hold.

  1. 1Grok 4.20xAI2,000,000 tokens
  2. 2Grok 4.20 Multi-AgentxAI2,000,000 tokens
  3. 3GPT-5.4OpenAI1,050,000 tokens
  4. 4GPT-5.4 ProOpenAI1,050,000 tokens
  5. 5GPT-5.5OpenAI1,050,000 tokens

Cheapest with a 1M-token context

Lowest input price among models that accept at least 1,000,000 tokens.

  1. 1Qwen3.7 FlashQwen$0.03 in · $0.13 out
  2. 2Qwen3.5-FlashQwen$0.065 in · $0.26 out
  3. 3Laguna S 2.1Poolside$0.09 in · $0.18 out
  4. 4Claude Haiku 5.5Anthropic$0.10 in · $0.50 out
  5. 5Gemini 2.5 Flash LiteGoogle$0.10 in · $0.40 out

Cheapest with image input and tool calling

For agents that need to see screenshots and call functions.

  1. 1Ling 3.0 Flash VLinclusionAI$0.021 in · $0.0616 out
  2. 2Qwen3.7 FlashQwen$0.03 in · $0.13 out
  3. 3Gemma 3 12BGoogle$0.05 in · $0.15 out
  4. 4Gemma 3 4BGoogle$0.05 in · $0.10 out
  5. 5GPT-5 NanoOpenAI$0.05 in · $0.40 out

Newest releases

Date each model was first listed by our data source.

  1. 1Step 5 PreviewStepFun2026-10-08
  2. 2Claude Haiku 5.5Anthropic2026-10-07
  3. 3Mistral Large 4Mistral2026-10-06
  4. 4Pareto 26.10 Previewunbiased2026-10-01
  5. 5GPT-6.1 SolOpenAI2026-09-29

Lists computed from the data on 2026-10-09. Sources: OpenRouter models API and LiteLLM model prices and context windows.

FAQ

Frequently asked questions

What does “Fill context” mean?

It is the price of one request whose prompt uses the model's entire context window, at the list input price (and the long-context price where one applies). It shows how expensive very long prompts get: across the models in the table it ranges from $0.000328 to $63.00 per request.

What is the difference between context window and max output?

The context window is the total number of tokens a request can hold, counting both your prompt and the reply. Max output is the most the model will generate in one reply. A model with a 1,000,000-token window and 65,536 max output can read a huge document but can only write about 50,000 words back in one go.

Why do input and output have different prices?

Reading your prompt can be done in parallel, but generating a reply happens one token at a time, which costs the provider more compute per token. That's why output is usually several times more expensive than input. To see what a real workload costs, use the LLM cost calculator.

What does “open weights” mean?

The model's weights are published, so you can download it and run it on your own hardware or pick from several hosting companies. The price shown for open-weight models is a typical hosted API price; running them yourself costs whatever your hardware and electricity cost.

How current is this table?

Prices, context windows and capabilities are refreshed automatically every day from the OpenRouter models API and the open-source LiteLLM price list, with manual corrections when a source is wrong. This table was last updated on 2026-10-09. See the methodology.

Is the cheapest model the best choice?

Not necessarily. Price says nothing about quality, speed or reliability for your task. Use the filters to shortlist models that meet your hard requirements (context size, image input, tool calling), then test the shortlist on your own prompts before choosing.