Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Tokens & Costs

AI model pricing and specs

A page for each of 20 widely used AI models: list prices, context window, output limit, what typical workloads cost, and cheaper alternatives with the same essentials.

Looking for a model that isn’t here? The comparison table covers all 340+ models, and the provider pages list each company’s full line-up.

Prices updated 2026-10-09 · USD per 1M tokens

About

How to read these numbers

Prices are in US dollars per million tokens, the unit every AI API bills in. Input is what you send: your prompt, the conversation so far, any documents. Output is what the model writes back. Across the 20 models here, output costs a median of 4.3 times as much as input, so a task that writes long answers costs far more than one that reads a long document and replies briefly.

The context window is the most tokens a single request can hold, prompt and reply together. The output limit is a separate, smaller cap on the reply: for the median model here that publishes one it is 13% of the window, and 6 publish no separate cap at all. A model with a huge window still stops writing when it reaches its output limit.

These are list prices. 18 of the 20 models also list a cheaper rate for cached input (repeated prompt prefixes), and some providers offer batch discounts or charge more above a prompt-length threshold; each model page shows what applies. To turn prices into a monthly bill for your own workload, use the cost calculator.

This directory covers 20 current, widely used models from the major providers, and a model only gets a page when its price, context window, output limit and release date are all known. Every other priced model is in the comparison table.

Further reading: Claude vs GPT vs Gemini pricing, explained with live prices, The cheapest LLM APIs right now, ranked from daily price data, Context windows explained: what counts, what happens at the limit, and how to check fit.

FAQ

Questions people ask

Where do these prices come from?

From the OpenRouter models API and LiteLLM’s open-source price list, checked against each provider’s own pricing pages, with manual corrections where the automated sources are wrong. The methodology page explains the process.

How often are the prices updated?

Every day. A scheduled job downloads the sources, validates every record and republishes the site. The prices on this page are from 2026-10-09.

Why does output cost more than input?

Providers set prices, and almost all of them charge more for output. The usual reason given is the work involved: a model reads your whole prompt in one pass, but writes its reply one token at a time, and each new token needs another pass through the model.

Which model here is the cheapest?

On input price, gpt-oss-120b at $0.037 per million tokens (data from 2026-10-09). The cheapest model for a real task also depends on output price, how many tokens it turns your text into and how long its answers are, so compare with your own numbers in the cost calculator.

Which model has the largest context window?

GPT-6.1 Sol, with 1,050,000 tokens. A bigger window lets you send more at once, but every token in it is billed, so it is worth checking how much of a document you really need to send. The context window checker shows which models fit your text.