Guide · Models
Claude vs GPT vs Gemini pricing, explained with live prices
Anthropic, OpenAI and Google each sell a top, middle and small model line, billed per million input and output tokens, with output costing 5× to 8.3× the input rate. What you actually pay depends on more than the list price: caching, batch discounts, long-prompt surcharges, thinking tokens billed as output, and how many tokens each tokenizer makes of your text.
By Tahir NazirUpdated 11 min read
On this page
How do Claude, GPT and Gemini prices compare?
Each provider sells a line of models at very different prices, so the useful comparison is tier by tier. Here are the standard API rates for each provider’s top, middle and small model as of 2026-10-09, in US dollars per million tokens:
Top tier
- Claude Fable 5.1Anthropic
- $10 per 1M input tokens$50 per 1M output tokens
- GPT-6 AstraOpenAI
- $10 per 1M input tokens$50 per 1M output tokens
- Gemini 3.1 Pro PreviewGoogle
- $2 per 1M input tokens$12 per 1M output tokens
Middle tier
- Claude Sonnet 5.5Anthropic
- $2 per 1M input tokens$10 per 1M output tokens
- GPT-6.1 SolOpenAI
- $2 per 1M input tokens$10 per 1M output tokens
- Gemini 3.8 FlashGoogle
- $0.75 per 1M input tokens$3.75 per 1M output tokens
Small tier
- Claude Haiku 5.5Anthropic
- $0.10 per 1M input tokens$0.50 per 1M output tokens
- GPT-6 LunaOpenAI
- $0.10 per 1M input tokens$0.50 per 1M output tokens
- Gemini 3.5 Flash LiteGoogle
- $0.30 per 1M input tokens$2.50 per 1M output tokens
USD per 1M tokens · standard rates · 2026-10-09 · bars scaled within each tier
How we picked them: the most capable model each provider sells, its mid-priced line (Sonnet, Sol and Flash) and its cheapest current line, going by the providers’ own model pages. Anthropic sells a fourth line, Opus, between Fable and Sonnet, and its docs suggest Claude Opus 5.5 as the default starting point for most workloads, so it’s in the full table. Google’s newest Pro model is still a preview.
| Model | Input / output | Cached input | Batch input / output | Long prompts | Context window |
|---|---|---|---|---|---|
| Claude Fable 5.1 | $10 / $50 | $0.25 | $5 / $25 | Same rate across the window | 1,000,000 tokens |
| Claude Opus 5.5 | $4 / $20 | $0.20 | $2 / $10 | Same rate across the window | 1,000,000 tokens |
| Claude Sonnet 5.5 | $2 / $10 | $0.10 | $1 / $5 | Same rate across the window | 1,000,000 tokens |
| Claude Haiku 5.5 | $0.10 / $0.50 | $0.01 | $0.05 / $0.25 | $0.50 / $2.50 over 100,000 prompt tokens | 1,000,000 tokens |
| GPT-6 Astra | $10 / $50 | $1 | $5 / $25 | $20 / $75 over 272,000 prompt tokens | 1,050,000 tokens |
| GPT-6.1 Sol | $2 / $10 | $0.10 | $1 / $5 | $4 / $15 over 272,000 prompt tokens | 1,050,000 tokens |
| GPT-6 Luna | $0.10 / $0.50 | $0.01 | $0.05 / $0.25 | $0.20 / $0.75 over 272,000 prompt tokens | 1,050,000 tokens |
| Gemini 3.1 Pro Preview | $2 / $12 | $0.20 | $1 / $6 | $4 / $18 over 200,000 prompt tokens | 1,048,576 tokens |
| Gemini 3.8 Flash | $0.75 / $3.75 | $0.075 | $0.375 / $1.875 | Same rate across the window | 1,048,576 tokens |
| Gemini 3.5 Flash Lite | $0.30 / $2.50 | $0.03 | $0.15 / $1.25 | Same rate across the window | 1,048,576 tokens |
From our daily data (OpenRouter models API and LiteLLM price list). Cached input is the price of reading from the prompt cache; writing to it can cost extra (see below).
On list price alone, the lowest input rate among the nine tier models belongs to Claude Haiku 5.5 and GPT-6 Luna, and the lowest output rate to Claude Haiku 5.5 and GPT-6 Luna. But list prices are only the starting point: the rest of this guide covers the billing rules that move the real number.
How does each provider bill its API?
All three charge separately for input tokens (everything you send) and output tokens (everything the model writes). The differences are in caching, discounts for slower processing, and surcharges for long prompts or special endpoints.
Anthropic (Claude)
- Caching is opt-in. You mark the reusable part of the prompt with
cache_control, or add one top-level field for automatic caching. Writing to the 5-minute cache costs 1.25× the input price and to the 1-hour cache 2×. Reads cost 2.5% to 10% of the input price on the current models. Prompts under 512 tokens aren’t cached on the current models. Our guide to prompt caching covers when it pays off. - Batch API: 50% off input and output for requests that can wait.
- Long prompts: Fable, Opus and Sonnet charge the same rate up to their full 1,000,000-token window. Claude Haiku 5.5 switches to $0.50 / $2.50 once a prompt passes 100,000 tokens.
- Extras: US-only inference (
inference_geo) costs 1.1×, and a faster “fast mode” on Opus is sold at a premium. Any request with tools also carries a hidden tool-use system prompt (286 tokens on Opus 5.5, Sonnet 5.5 and Haiku 5.5 with the defaulttool_choice), billed as input.
OpenAI (GPT)
- Caching is automatic on supported models once a prompt reaches 1,024 tokens (GPT-5.6 and later), and a cached prefix stays reusable for 30 minutes after its last use. On those models a cache write costs 1.25× the input price; reads cost 5% to 10% of input on the three GPT-6 models here.
- Processing tiers: Batch is half price and Flex is also discounted; Fast and Ultrafast cost more.
- Long prompts: above 272,000 input tokens the whole request moves to a long-context rate. For GPT-6.1 Sol that’s $4 / $15.
- Regional processing (data residency endpoints) adds 10% for models released on or after 5 March 2026.
Google (Gemini)
- Free tier: most Gemini models can be used free within usage limits, which is handy for prototyping. See our guide to getting a free Gemini API key.
- Caching is automatic (implicit) on Gemini 2.5 and newer, with a minimum of 4,096 tokens on the current Flash and Pro models. Reads cost 10% of input on the three models here. Explicit caches you create yourself also carry an hourly storage charge.
- Batch and Flex are half price; Priority costs more.
- Long prompts: Pro models charge more once a prompt passes 200,000 tokens. For Gemini 3.1 Pro Preview that’s $4 / $18.
- Scheduled change: Google’s pricing page lists doubled prices for Gemini 3.6, 3.7 and 3.8 Flash from 1 January 2027. Our data refreshes daily, so the tables on this page will follow.
Do thinking tokens cost extra?
Yes. When a model reasons before it answers, those thinking tokens are billed at the output price on all three APIs, and you usually don’t see them. OpenAI says reasoning tokens are “billed as output tokens” and aren’t visible through the API. Anthropic bills the full thinking, even when it returns only a summary or nothing at all. Google’s output price “includes thinking tokens”.
Because output is the expensive side, thinking can matter more than the input price. In the worked example below, adding 1,000 thinking tokens to each answer raises the monthly bill by between 143% and 172%, depending on the model. Thinking tokens also count towards the output limit you set (max_tokens on Claude, max_output_tokens on OpenAI), so a low limit can cut an answer short.
Each provider lets you turn the amount of thinking up or down: effort on Claude, reasoning.effort on OpenAI and thinking_level on Gemini 3. Lower settings are cheaper and faster; test whether your task still comes out right.
Why per-token prices aren’t directly comparable
A price per million tokens only compares fairly if your text becomes the same number of tokens on each model, and it doesn’t. Every model family cuts text into tokens with its own tokenizer, so the same prompt can be noticeably more tokens on one model than another. Our guide to what a token is shows how much this varies.
- Even within one family. Anthropic says Claude Opus 4.7 and later use a newer tokenizer that produces approximately 30% more tokens for the same text than the one before. A model with the same per-token price as its predecessor can cost more per request.
- Few public tokenizers for the current models. OpenAI’s tiktoken library maps GPT-5 and earlier to published encodings, but no GPT-6 model (as of version 0.14.0). Anthropic doesn’t publish its tokenizer; Anthropic’s
count_tokensand Google’scountTokensendpoints give the official counts. - Hidden input. Tool definitions, tool-use system prompts, images and earlier turns of a conversation are all billed as input, and each provider formats them differently.
So when two models are listed at exactly the same rates, as Claude Sonnet 5.5 and GPT-6.1 Sol are today, the one that turns your text into fewer tokens, and answers in fewer, is the cheaper one. The only way to know is to count your own prompts on each.
Free toolAI token counterCount the same text on GPT, Claude and Gemini side by side, with official Claude and Gemini counts when you add your own key.Worked example: a support assistant answering 1,000 questions a day
Say each request sends a 5,000-token fixed prefix (system prompt and help-centre extract) plus 1,000 tokens of question and history, and gets a 400-token answer, 1,000 times a day. The prefix is long enough to be cached on all three. We use the same token counts for every model so that only the prices differ:
| Model | No caching | Prefix cached | Cached + 1,000 thinking tokens |
|---|---|---|---|
| Claude Haiku 5.5 | $24.33 | $10.65 | $25.85 |
| GPT-6 Luna | $24.33 | $10.65 | $25.85 |
| Gemini 3.5 Flash Lite | $85.17 | $44.10 | $120.15 |
| Gemini 3.8 Flash | $182.50 | $79.84 | $193.91 |
| Claude Sonnet 5.5 | $486.67 | $197.71 | $501.88 |
| GPT-6.1 Sol | $486.67 | $197.71 | $501.88 |
| Gemini 3.1 Pro Preview | $511.00 | $237.25 | $602.25 |
| Claude Opus 5.5 | $973.33 | $395.42 | $1,004 |
| Claude Fable 5.1 | $2,433 | $950.52 | $2,471 |
| GPT-6 Astra | $2,433 | $1,065 | $2,585 |
1,000 requests a day over an average month (30.4 days). Assumes a warm cache and ignores cache-write charges, which add a little on Claude and on newer GPT models.
With caching on, the cheapest here are Claude Haiku 5.5 and GPT-6 Luna at $10.65 a month and the most expensive is GPT-6 Astra at $1,065, about 100 times as much. Caching the prefix cuts each bill by between 48% and 61%, which is why it’s worth setting up before you compare anything else.
What changes with one very long prompt
Long-prompt surcharges only bite on big requests. Here is a single request with a 300,000-token prompt (a long contract or a large slice of a codebase) and a 1,000-token answer:
| Model | Cost of the request | Long-prompt rate applied? |
|---|---|---|
| Claude Sonnet 5.5 | $0.61 | No |
| GPT-6.1 Sol | $1.21 | Yes, over 272,000 tokens |
| Gemini 3.1 Pro Preview | $1.22 | Yes, over 200,000 tokens |
| Gemini 3.8 Flash | $0.23 | No |
| Claude Haiku 5.5 | $0.15 | Yes, over 100,000 tokens |
| GPT-6 Luna | $0.0608 | Yes, over 272,000 tokens |
When a long-prompt rate applies, it covers the whole request, not just the tokens above the threshold.
How to choose between Claude, GPT and Gemini on cost
Start with quality, then price the models that pass. A cheaper model that needs two attempts, or a longer prompt, isn’t cheaper.
- Pick the tier the task needs. Test a small model first for classification, extraction and routing; move up only when your own examples fail.
- Count your real prompts on each candidate with the token counter, because tokenizers differ.
- Price the real shape of the workload: how much of the prompt repeats (caching), whether it can wait (batch), how long answers and thinking run.
- Check the thresholds. If prompts can pass 100,000 tokens, see which models charge more for long prompts, and whether a summary or retrieval would avoid it.
- Check where you can buy it. Claude is also sold through Amazon Bedrock, Google Cloud and Microsoft Foundry. Bedrock and Google Cloud set their own prices, and their regional endpoints cost 10% more than global ones for recent Claude models.
- Recheck regularly. Prices on this page change with the daily data; the model comparison table shows every model’s current rates.
FAQ
Questions people ask
Is Claude more expensive than GPT?
Compare tier by tier. Per million input / output tokens as of 2026-10-09: Claude Fable 5.1 and GPT-6 Astra both cost $10 / $50; Claude Sonnet 5.5 and GPT-6.1 Sol both cost $2 / $10; Claude Haiku 5.5 and GPT-6 Luna both cost $0.10 / $0.50. Caching rates, long-prompt surcharges and tokenizers also differ, so the same task can cost more on one even when list prices match. Count your own text before deciding.
Which is cheapest: Gemini, Claude or GPT?
There’s no single answer, because each sells cheap and expensive models. In our support-assistant example with caching, Claude Haiku 5.5 and GPT-6 Luna came out cheapest at $10.65 a month. Free tiers, batch discounts, long-prompt surcharges and tokenizer differences can all change the order for your workload, so price your own numbers in a cost calculator.
Do reasoning or thinking tokens cost extra?
They are billed as output tokens on Claude, GPT and Gemini, at the normal output price, even though you usually can’t see them. A model that thinks for 1,000 tokens before a 400-token answer bills 1,400 output tokens. Lower the effort or thinking level when a task doesn’t need deep reasoning.
Is the Gemini API free?
Partly. Google’s pricing page lists a free tier for most Gemini models, with usage limits; Gemini 3.1 Pro Preview is paid only. Paid use is billed per token like the others. Anthropic gives new API accounts a small amount of free credit to test with, then bills per token.
Are Claude prices the same on Amazon Bedrock and Google Cloud?
Not necessarily. Anthropic says Bedrock and Google Cloud set and invoice their own prices, so check their pricing pages. For Claude Sonnet 4.5, Haiku 4.5, Opus 4.5 and later models, their regional and multi-region endpoints cost 10% more than global endpoints.
Does prompt caching work the same way on all three?
No. OpenAI and Gemini cache repeated prompt prefixes automatically once they pass a minimum length. On Claude you turn caching on with cache_control, and writing to the cache costs more than normal input, while reads cost a small fraction of it. All three need the repeated part at the start of the prompt.
Try it
Tools from this guide
Keep reading