Tokens & Costs
AI model comparison: prices, context windows and features of 340+ LLMs
Compare every major AI model's API price, context window, output limit and capabilities in one sortable table, updated daily.
Showing 50 of 333 matching models (333 in total)
| Cost estimate | ||||||||
|---|---|---|---|---|---|---|---|---|
| Step 5 PreviewStepFun | $1 | $2.70 | $0.05 | 1M | 64K | $1.00 | 2026-10-08 | Estimate cost |
| Claude Haiku 5.5Anthropic | $0.10 | $0.50 | $0.01 | 1M | 128K | $0.50 | 2026-10-07 | Estimate cost |
| Mistral Large 4Mistral | $0.68 | $2.09 | $0.07 | 1.05M | 262K | $0.71 | 2026-10-06 | Estimate cost |
| Ling 3.1 FlashinclusionAI | $0 | $0 | – | 262K | 33K | $0.00 | 2026-10-02 | Estimate cost |
| Pareto 26.10 Previewunbiased | $0.80 | $3.20 | $0.03 | 1.05M | 131K | $0.84 | 2026-10-01 | Estimate cost |
| GPT-6.1 SolOpenAI | $2 | $10 | $0.10 | 1.05M | 128K | $4.20 | 2026-09-29 | Estimate cost |
| GPT-6.1 Sol ProOpenAI | $2 | $10 | $0.10 | 1.05M | 128K | $4.20 | 2026-09-29 | Estimate cost |
| Claude Sonnet 5.5Anthropic | $2 | $10 | $0.10 | 1M | 128K | $2.00 | 2026-09-28 | Estimate cost |
| Perceptron Mk1.5Perceptron | $0.15 | $1.50 | – | 37K | 8K | $0.00553 | 2026-09-25 | Estimate cost |
| Ember-1Fireworks | $3 | $15 | $0.30 | 1.05M | – | $3.15 | 2026-09-24 | Estimate cost |
| Aion 3.5AionLabs | $3 | $6 | $0.75 | 262K | 33K | $0.79 | 2026-09-23 | Estimate cost |
| Aion 3.5 MiniAionLabs | $0.70 | $1.40 | $0.18 | 262K | 33K | $0.18 | 2026-09-23 | Estimate cost |
| GLM 5.3 PrimeZ.ai | $2.80 | $8.80 | $0.56 | 1M | 131K | $2.80 | 2026-09-23 | Estimate cost |
| Qwen3.8 Max PrimeQwen | $4 | $12 | $0.50 | 1M | 131K | $4.00 | 2026-09-23 | Estimate cost |
| Solar Mini 4Upstage | $0.05 | $0.20 | $0.005 | 524K | 131K | $0.0262 | 2026-09-23 | Estimate cost |
| Claude Opus 5.5Anthropic | $4 | $20 | $0.20 | 1M | 128K | $4.00 | 2026-09-22 | Estimate cost |
| Command A+Cohere | $0.30 | $1.50 | $0.15 | 192K | 64K | $0.0576 | 2026-09-22 | Estimate cost |
| GPT-6 LunaOpenAI | $0.10 | $0.50 | $0.01 | 1.05M | 128K | $0.21 | 2026-09-22 | Estimate cost |
| GPT-6 Luna ProOpenAI | $0.10 | $0.50 | $0.01 | 1.05M | 128K | $0.21 | 2026-09-22 | Estimate cost |
| GPT-6 SolOpenAI | $2 | $10 | $0.20 | 1.05M | 128K | $4.20 | 2026-09-22 | Estimate cost |
| GPT-6 Sol ProOpenAI | $2 | $10 | $0.20 | 1.05M | 128K | $4.20 | 2026-09-22 | Estimate cost |
| Grok 4.7xAI | $2 | $6 | $0.50 | 500K | – | $2.00 | 2026-09-21 | Estimate cost |
| MiMo-V2.6-FlashXiaomiOpen | $0.14 | $0.28 | $0.0028 | 1.05M | 131K | $0.15 | 2026-09-21 | Estimate cost |
| MiMo-V2.6-ProXiaomiOpen | $0.435 | $0.87 | $0.0036 | 1.05M | 131K | $0.46 | 2026-09-21 | Estimate cost |
| MiMo-V2.6-Pro-UltraSpeedXiaomi | $4.35 | $8.70 | $0.036 | 1.05M | 131K | $4.56 | 2026-09-21 | Estimate cost |
| Qwen3.8 Omni FlashQwen | $0.15 | $0.47 | $0.016 | 1M | 131K | $0.15 | 2026-09-21 | Estimate cost |
| GLM 5.3 FlashXZ.ai | $0.37 | $1.25 | $0.09 | 1.05M | 131K | $0.39 | 2026-09-18 | Estimate cost |
| Ternary Bonsai 2 27BPrismMLOpen | $0.075 | $0.50 | $0.0375 | 262K | 33K | $0.0197 | 2026-09-18 | Estimate cost |
| Paretounbiased | $2.50 | $7.50 | $0.25 | 262K | 131K | $0.66 | 2026-09-17 | Estimate cost |
| Schematron V2 SmallInference.netOpen | $0.05 | $0.23 | $0.05 | 128K | 4K | $0.0064 | 2026-09-12 | Estimate cost |
| Schematron V2 TurboInference.netOpen | $0.03 | $0.15 | $0.03 | 128K | 8K | $0.00384 | 2026-09-12 | Estimate cost |
| Fugu MaxSakana | $2 | $6 | $0.25 | 1M | 128K | $2.00 | 2026-09-11 | Estimate cost |
| Fugu Ultra v2Sakana | $5 | $30 | $0.50 | 1M | 128K | $10.00 | 2026-09-11 | Estimate cost |
| DeepSeek V4.1 FlashDeepSeekOpen | $0.30 | $1.20 | $0.006 | 1.05M | – | $0.31 | 2026-09-10 | Estimate cost |
| Ling 3.0 Flash VLinclusionAIOpen | $0.021 | $0.0616 | $0.0042 | 262K | 33K | $0.00551 | 2026-09-10 | Estimate cost |
| Mercury 2.5Inception | $0.04 | $0.15 | $0.004 | 260K | 66K | $0.0104 | 2026-09-08 | Estimate cost |
| Nex-N2.5-MiniNex AGIOpen | $0.025 | $0.10 | $0.0025 | 262K | – | $0.00655 | 2026-09-08 | Estimate cost |
| Nex-N2.5-ProNex AGIOpen | $0.075 | $0.25 | $0.015 | 262K | – | $0.0197 | 2026-09-08 | Estimate cost |
| GPT-6 AstraOpenAI | $10 | $50 | $1 | 1.05M | 128K | $21.00 | 2026-09-04 | Estimate cost |
| GPT-6 Astra ProOpenAI | $10 | $50 | $1 | 1.05M | 128K | $21.00 | 2026-09-04 | Estimate cost |
| Ling 3.0 Flash SanteinclusionAI | $0.042 | $0.1232 | $0.0084 | 262K | 33K | $0.011 | 2026-09-04 | Estimate cost |
| Qwen3.8 Max (0902)Qwen | $2 | $6 | $0.25 | 1M | 131K | $2.00 | 2026-09-03 | Estimate cost |
| Gemini 3.8 FlashGoogle | $0.75 | $3.75 | $0.075 | 1.05M | 66K | $0.79 | 2026-09-02 | Estimate cost |
| Muse Spark 1.3Meta | $1.25 | $4.25 | $0.15 | 1.05M | – | $1.31 | 2026-09-02 | Estimate cost |
| Muse Spark 1.3 ContributorMeta | $0.10 | $0.20 | $0.002 | 1.05M | – | $0.10 | 2026-09-02 | Estimate cost |
| Claude Fable 5.1Anthropic | $10 | $50 | $0.25 | 1M | 128K | $10.00 | 2026-09-01 | Estimate cost |
| Granite 4.2 8BIBMOpen | $0.06 | $0.25 | $0.015 | 131K | – | $0.00786 | 2026-08-31 | Estimate cost |
| Ling 3.0 Flash FininclusionAI | $0.042 | $0.1232 | $0.0084 | 262K | 33K | $0.011 | 2026-08-27 | Estimate cost |
| GLM 5.3 FlashZ.aiOpen | $0.15 | $0.50 | $0.03 | 1.05M | – | $0.16 | 2026-08-26 | Estimate cost |
| Qwen3.8 FlashQwenOpen | $0.15 | $0.47 | $0.016 | 1M | 131K | $0.15 | 2026-08-26 | Estimate cost |
Steps
How to use the AI model comparison
- Search for a model, or pick one or more providers.
- Set a minimum context window and any capabilities you need, such as image input or tool calling.
- Click a column header to sort, for example by input price or context window. Click again to reverse.
- Use “Estimate cost” on any row to price your own workload in the LLM cost calculator.
- Copy the page address to share the exact filters and sort you’re looking at.
Method
How it works
Choosing a model usually starts with three questions: can it handle my input, can it do what I need, and what will it cost? This table answers all three for 340+ models at once, so you can narrow the field before testing anything.
Reading the prices
Prices are list prices in US dollars per million tokens. Input covers everything you send (system prompt, history, documents); output is the model’s reply and costs more because it’s generated one token at a time. Cached input is the discounted price for re-reading a prompt prefix the provider has already cached, which matters for chatbots and agents that resend the same instructions on every request. A dash means the provider hasn’t published that price.
Context window, max output and the cost of filling it
The context window is the most a single request can hold, prompt and reply together. Max output limits the reply on its own. The Fill context column multiplies the window by the input price, which tells you what one maximum-length prompt costs. Two models can share a 1M-token window and still differ by more than a hundred times on this column, so it’s a quick way to see which long-context models are practical for everyday use.
Capabilities and open weights
The capability filters (image input, tool calling, reasoning, prompt caching, structured output) use the features each provider declares for its API. They tell you a feature exists, not how well it works. Open-weight models publish their weights, so several companies host them and you can also run them yourself; their price here is a typical hosted price, and our VRAM calculator covers self-hosting.
Where the data comes from
The table is generated from the OpenRouter models API and the open-source LiteLLM price list, refreshed every day. We correct the data by hand when a source is wrong; for example, when a feed reports a cache price as the normal input price, the model is held back until it’s checked against the provider’s pricing page. Release dates are when a model was first listed by our source, which can differ from the announcement date. The methodology page has the details.
Examples
Worked examples
Largest context windows
Most tokens a single request can hold.
- 1Grok 4.20xAI2,000,000 tokens
- 2Grok 4.20 Multi-AgentxAI2,000,000 tokens
- 3GPT-5.4OpenAI1,050,000 tokens
- 4GPT-5.4 ProOpenAI1,050,000 tokens
- 5GPT-5.5OpenAI1,050,000 tokens
Cheapest with a 1M-token context
Lowest input price among models that accept at least 1,000,000 tokens.
- 1Qwen3.7 FlashQwen$0.03 in · $0.13 out
- 2Qwen3.5-FlashQwen$0.065 in · $0.26 out
- 3Laguna S 2.1Poolside$0.09 in · $0.18 out
- 4Claude Haiku 5.5Anthropic$0.10 in · $0.50 out
- 5Gemini 2.5 Flash LiteGoogle$0.10 in · $0.40 out
Cheapest with image input and tool calling
For agents that need to see screenshots and call functions.
- 1Ling 3.0 Flash VLinclusionAI$0.021 in · $0.0616 out
- 2Qwen3.7 FlashQwen$0.03 in · $0.13 out
- 3Gemma 3 12BGoogle$0.05 in · $0.15 out
- 4Gemma 3 4BGoogle$0.05 in · $0.10 out
- 5GPT-5 NanoOpenAI$0.05 in · $0.40 out
Newest releases
Date each model was first listed by our data source.
- 1Step 5 PreviewStepFun2026-10-08
- 2Claude Haiku 5.5Anthropic2026-10-07
- 3Mistral Large 4Mistral2026-10-06
- 4Pareto 26.10 Previewunbiased2026-10-01
- 5GPT-6.1 SolOpenAI2026-09-29
Lists computed from the data on 2026-10-09. Sources: OpenRouter models API and LiteLLM model prices and context windows.
FAQ
Frequently asked questions
What does “Fill context” mean?
It is the price of one request whose prompt uses the model's entire context window, at the list input price (and the long-context price where one applies). It shows how expensive very long prompts get: across the models in the table it ranges from $0.000328 to $63.00 per request.
What is the difference between context window and max output?
The context window is the total number of tokens a request can hold, counting both your prompt and the reply. Max output is the most the model will generate in one reply. A model with a 1,000,000-token window and 65,536 max output can read a huge document but can only write about 50,000 words back in one go.
Why do input and output have different prices?
Reading your prompt can be done in parallel, but generating a reply happens one token at a time, which costs the provider more compute per token. That's why output is usually several times more expensive than input. To see what a real workload costs, use the LLM cost calculator.
What does “open weights” mean?
The model's weights are published, so you can download it and run it on your own hardware or pick from several hosting companies. The price shown for open-weight models is a typical hosted API price; running them yourself costs whatever your hardware and electricity cost.
How current is this table?
Prices, context windows and capabilities are refreshed automatically every day from the OpenRouter models API and the open-source LiteLLM price list, with manual corrections when a source is wrong. This table was last updated on 2026-10-09. See the methodology.
Is the cheapest model the best choice?
Not necessarily. Price says nothing about quality, speed or reliability for your task. Use the filters to shortlist models that meet your hard requirements (context size, image input, tool calling), then test the shortlist on your own prompts before choosing.
Related
Related tools
- AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.
- Context Window CheckerSee whether your text fits each model's context window.
- Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.
- GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.
- Claude Code Error DatabaseExact Claude Code error messages with tested fixes.
- CLAUDE.md GeneratorGenerate a CLAUDE.md and AGENTS.md for your project.
- MCP Config GeneratorBuild MCP server configs for Claude, Cursor, VS Code and Devin Desktop (Windsurf).
- Gemini API PlaygroundTry the Gemini API free with your own key.
- AI Crawler robots.txt GeneratorBlock or allow GPTBot, ClaudeBot and other AI crawlers.
- JSON RepairFix broken JSON from LLM output.
- Free Gemini API Key GuideGet a free Gemini API key in a few minutes.