Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Guide · Models

Claude vs GPT vs Gemini pricing, explained with live prices

Anthropic, OpenAI and Google each sell a top, middle and small model line, billed per million input and output tokens, with output costing 5× to 8.3× the input rate. What you actually pay depends on more than the list price: caching, batch discounts, long-prompt surcharges, thinking tokens billed as output, and how many tokens each tokenizer makes of your text.

By Tahir NazirUpdated 11 min read

On this page
  1. How do Claude, GPT and Gemini prices compare?
  2. How does each provider bill its API?
  3. Do thinking tokens cost extra?
  4. Why per-token prices aren’t directly comparable
  5. Worked example: a support assistant answering 1,000 questions a day
  6. How to choose between Claude, GPT and Gemini on cost
  7. Questions people ask

How do Claude, GPT and Gemini prices compare?

Each provider sells a line of models at very different prices, so the useful comparison is tier by tier. Here are the standard API rates for each provider’s top, middle and small model as of 2026-10-09, in US dollars per million tokens:

Bars are scaled within each tier, because the tiers differ too much to share one axis. Compare the three models in a tier with each other, not across tiers.

How we picked them: the most capable model each provider sells, its mid-priced line (Sonnet, Sol and Flash) and its cheapest current line, going by the providers’ own model pages. Anthropic sells a fourth line, Opus, between Fable and Sonnet, and its docs suggest Claude Opus 5.5 as the default starting point for most workloads, so it’s in the full table. Google’s newest Pro model is still a preview.

Current lineups, standard API rates (USD per 1M tokens, 2026-10-09)
ModelInput / outputCached inputBatch input / outputLong promptsContext window
Claude Fable 5.1$10 / $50$0.25$5 / $25Same rate across the window1,000,000 tokens
Claude Opus 5.5$4 / $20$0.20$2 / $10Same rate across the window1,000,000 tokens
Claude Sonnet 5.5$2 / $10$0.10$1 / $5Same rate across the window1,000,000 tokens
Claude Haiku 5.5$0.10 / $0.50$0.01$0.05 / $0.25$0.50 / $2.50 over 100,000 prompt tokens1,000,000 tokens
GPT-6 Astra$10 / $50$1$5 / $25$20 / $75 over 272,000 prompt tokens1,050,000 tokens
GPT-6.1 Sol$2 / $10$0.10$1 / $5$4 / $15 over 272,000 prompt tokens1,050,000 tokens
GPT-6 Luna$0.10 / $0.50$0.01$0.05 / $0.25$0.20 / $0.75 over 272,000 prompt tokens1,050,000 tokens
Gemini 3.1 Pro Preview$2 / $12$0.20$1 / $6$4 / $18 over 200,000 prompt tokens1,048,576 tokens
Gemini 3.8 Flash$0.75 / $3.75$0.075$0.375 / $1.875Same rate across the window1,048,576 tokens
Gemini 3.5 Flash Lite$0.30 / $2.50$0.03$0.15 / $1.25Same rate across the window1,048,576 tokens

From our daily data (OpenRouter models API and LiteLLM price list). Cached input is the price of reading from the prompt cache; writing to it can cost extra (see below).

On list price alone, the lowest input rate among the nine tier models belongs to Claude Haiku 5.5 and GPT-6 Luna, and the lowest output rate to Claude Haiku 5.5 and GPT-6 Luna. But list prices are only the starting point: the rest of this guide covers the billing rules that move the real number.

How does each provider bill its API?

All three charge separately for input tokens (everything you send) and output tokens (everything the model writes). The differences are in caching, discounts for slower processing, and surcharges for long prompts or special endpoints.

Anthropic (Claude)

  • Caching is opt-in. You mark the reusable part of the prompt with cache_control, or add one top-level field for automatic caching. Writing to the 5-minute cache costs 1.25× the input price and to the 1-hour cache 2×. Reads cost 2.5% to 10% of the input price on the current models. Prompts under 512 tokens aren’t cached on the current models. Our guide to prompt caching covers when it pays off.
  • Batch API: 50% off input and output for requests that can wait.
  • Long prompts: Fable, Opus and Sonnet charge the same rate up to their full 1,000,000-token window. Claude Haiku 5.5 switches to $0.50 / $2.50 once a prompt passes 100,000 tokens.
  • Extras: US-only inference (inference_geo) costs 1.1×, and a faster “fast mode” on Opus is sold at a premium. Any request with tools also carries a hidden tool-use system prompt (286 tokens on Opus 5.5, Sonnet 5.5 and Haiku 5.5 with the default tool_choice), billed as input.

OpenAI (GPT)

  • Caching is automatic on supported models once a prompt reaches 1,024 tokens (GPT-5.6 and later), and a cached prefix stays reusable for 30 minutes after its last use. On those models a cache write costs 1.25× the input price; reads cost 5% to 10% of input on the three GPT-6 models here.
  • Processing tiers: Batch is half price and Flex is also discounted; Fast and Ultrafast cost more.
  • Long prompts: above 272,000 input tokens the whole request moves to a long-context rate. For GPT-6.1 Sol that’s $4 / $15.
  • Regional processing (data residency endpoints) adds 10% for models released on or after 5 March 2026.

Google (Gemini)

  • Free tier: most Gemini models can be used free within usage limits, which is handy for prototyping. See our guide to getting a free Gemini API key.
  • Caching is automatic (implicit) on Gemini 2.5 and newer, with a minimum of 4,096 tokens on the current Flash and Pro models. Reads cost 10% of input on the three models here. Explicit caches you create yourself also carry an hourly storage charge.
  • Batch and Flex are half price; Priority costs more.
  • Long prompts: Pro models charge more once a prompt passes 200,000 tokens. For Gemini 3.1 Pro Preview that’s $4 / $18.
  • Scheduled change: Google’s pricing page lists doubled prices for Gemini 3.6, 3.7 and 3.8 Flash from 1 January 2027. Our data refreshes daily, so the tables on this page will follow.

Do thinking tokens cost extra?

Yes. When a model reasons before it answers, those thinking tokens are billed at the output price on all three APIs, and you usually don’t see them. OpenAI says reasoning tokens are “billed as output tokens” and aren’t visible through the API. Anthropic bills the full thinking, even when it returns only a summary or nothing at all. Google’s output price “includes thinking tokens”.

Because output is the expensive side, thinking can matter more than the input price. In the worked example below, adding 1,000 thinking tokens to each answer raises the monthly bill by between 143% and 172%, depending on the model. Thinking tokens also count towards the output limit you set (max_tokens on Claude, max_output_tokens on OpenAI), so a low limit can cut an answer short.

Each provider lets you turn the amount of thinking up or down: effort on Claude, reasoning.effort on OpenAI and thinking_level on Gemini 3. Lower settings are cheaper and faster; test whether your task still comes out right.

Why per-token prices aren’t directly comparable

A price per million tokens only compares fairly if your text becomes the same number of tokens on each model, and it doesn’t. Every model family cuts text into tokens with its own tokenizer, so the same prompt can be noticeably more tokens on one model than another. Our guide to what a token is shows how much this varies.

  • Even within one family. Anthropic says Claude Opus 4.7 and later use a newer tokenizer that produces approximately 30% more tokens for the same text than the one before. A model with the same per-token price as its predecessor can cost more per request.
  • Few public tokenizers for the current models. OpenAI’s tiktoken library maps GPT-5 and earlier to published encodings, but no GPT-6 model (as of version 0.14.0). Anthropic doesn’t publish its tokenizer; Anthropic’s count_tokens and Google’s countTokens endpoints give the official counts.
  • Hidden input. Tool definitions, tool-use system prompts, images and earlier turns of a conversation are all billed as input, and each provider formats them differently.

So when two models are listed at exactly the same rates, as Claude Sonnet 5.5 and GPT-6.1 Sol are today, the one that turns your text into fewer tokens, and answers in fewer, is the cheaper one. The only way to know is to count your own prompts on each.

Free toolAI token counterCount the same text on GPT, Claude and Gemini side by side, with official Claude and Gemini counts when you add your own key.

Worked example: a support assistant answering 1,000 questions a day

Say each request sends a 5,000-token fixed prefix (system prompt and help-centre extract) plus 1,000 tokens of question and history, and gets a 400-token answer, 1,000 times a day. The prefix is long enough to be cached on all three. We use the same token counts for every model so that only the prices differ:

Monthly cost of the support assistant, cheapest first (prices as of 2026-10-09)
ModelNo cachingPrefix cachedCached + 1,000 thinking tokens
Claude Haiku 5.5$24.33$10.65$25.85
GPT-6 Luna$24.33$10.65$25.85
Gemini 3.5 Flash Lite$85.17$44.10$120.15
Gemini 3.8 Flash$182.50$79.84$193.91
Claude Sonnet 5.5$486.67$197.71$501.88
GPT-6.1 Sol$486.67$197.71$501.88
Gemini 3.1 Pro Preview$511.00$237.25$602.25
Claude Opus 5.5$973.33$395.42$1,004
Claude Fable 5.1$2,433$950.52$2,471
GPT-6 Astra$2,433$1,065$2,585

1,000 requests a day over an average month (30.4 days). Assumes a warm cache and ignores cache-write charges, which add a little on Claude and on newer GPT models.

With caching on, the cheapest here are Claude Haiku 5.5 and GPT-6 Luna at $10.65 a month and the most expensive is GPT-6 Astra at $1,065, about 100 times as much. Caching the prefix cuts each bill by between 48% and 61%, which is why it’s worth setting up before you compare anything else.

What changes with one very long prompt

Long-prompt surcharges only bite on big requests. Here is a single request with a 300,000-token prompt (a long contract or a large slice of a codebase) and a 1,000-token answer:

One 300,000-token request, no caching
ModelCost of the requestLong-prompt rate applied?
Claude Sonnet 5.5$0.61No
GPT-6.1 Sol$1.21Yes, over 272,000 tokens
Gemini 3.1 Pro Preview$1.22Yes, over 200,000 tokens
Gemini 3.8 Flash$0.23No
Claude Haiku 5.5$0.15Yes, over 100,000 tokens
GPT-6 Luna$0.0608Yes, over 272,000 tokens

When a long-prompt rate applies, it covers the whole request, not just the tokens above the threshold.

How to choose between Claude, GPT and Gemini on cost

Start with quality, then price the models that pass. A cheaper model that needs two attempts, or a longer prompt, isn’t cheaper.

  1. Pick the tier the task needs. Test a small model first for classification, extraction and routing; move up only when your own examples fail.
  2. Count your real prompts on each candidate with the token counter, because tokenizers differ.
  3. Price the real shape of the workload: how much of the prompt repeats (caching), whether it can wait (batch), how long answers and thinking run.
  4. Check the thresholds. If prompts can pass 100,000 tokens, see which models charge more for long prompts, and whether a summary or retrieval would avoid it.
  5. Check where you can buy it. Claude is also sold through Amazon Bedrock, Google Cloud and Microsoft Foundry. Bedrock and Google Cloud set their own prices, and their regional endpoints cost 10% more than global ones for recent Claude models.
  6. Recheck regularly. Prices on this page change with the daily data; the model comparison table shows every model’s current rates.

FAQ

Questions people ask

Is Claude more expensive than GPT?

Compare tier by tier. Per million input / output tokens as of 2026-10-09: Claude Fable 5.1 and GPT-6 Astra both cost $10 / $50; Claude Sonnet 5.5 and GPT-6.1 Sol both cost $2 / $10; Claude Haiku 5.5 and GPT-6 Luna both cost $0.10 / $0.50. Caching rates, long-prompt surcharges and tokenizers also differ, so the same task can cost more on one even when list prices match. Count your own text before deciding.

Which is cheapest: Gemini, Claude or GPT?

There’s no single answer, because each sells cheap and expensive models. In our support-assistant example with caching, Claude Haiku 5.5 and GPT-6 Luna came out cheapest at $10.65 a month. Free tiers, batch discounts, long-prompt surcharges and tokenizer differences can all change the order for your workload, so price your own numbers in a cost calculator.

Do reasoning or thinking tokens cost extra?

They are billed as output tokens on Claude, GPT and Gemini, at the normal output price, even though you usually can’t see them. A model that thinks for 1,000 tokens before a 400-token answer bills 1,400 output tokens. Lower the effort or thinking level when a task doesn’t need deep reasoning.

Is the Gemini API free?

Partly. Google’s pricing page lists a free tier for most Gemini models, with usage limits; Gemini 3.1 Pro Preview is paid only. Paid use is billed per token like the others. Anthropic gives new API accounts a small amount of free credit to test with, then bills per token.

Are Claude prices the same on Amazon Bedrock and Google Cloud?

Not necessarily. Anthropic says Bedrock and Google Cloud set and invoice their own prices, so check their pricing pages. For Claude Sonnet 4.5, Haiku 4.5, Opus 4.5 and later models, their regional and multi-region endpoints cost 10% more than global endpoints.

Does prompt caching work the same way on all three?

No. OpenAI and Gemini cache repeated prompt prefixes automatically once they pass a minimum length. On Claude you turn caching on with cache_control, and writing to the cache costs more than normal input, while reads cost a small fraction of it. All three need the repeated part at the start of the prompt.

Try it

Tools from this guide

Keep reading