Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Guide · Tokens & costs

What is a token in AI? A plain-English guide with real examples

A token is the unit a language model reads and writes: usually a word, part of a word, or a punctuation mark. In English, one token averages about four characters or three quarters of a word, and every API price and context limit is measured in them.

By Tahir NazirUpdated 7 min read

On this page
  1. What a token is
  2. How many tokens is a word?
  3. Why other languages use more tokens
  4. Different models, different counts
  5. Why tokens matter: cost and context
  6. How to use fewer tokens
  7. Questions people ask

What a token is

Language models don’t see letters or words. Before your text reaches the model, a tokenizer cuts it into pieces from a fixed vocabulary, typically 100,000 to 260,000 entries, and replaces each piece with a number. The model reads that list of numbers and predicts the next number, one at a time. Those pieces are tokens.

Common words get a token of their own. Rarer words are built from several pieces, and a space is usually glued to the front of the word that follows it. Each token also has an ID, its position in the vocabulary, and that number is all the model actually sees:

One sentence, 10 tokens. “Tokenizers” becomes two pieces, “unbelievable” stays whole because it follows a space, and the year splits into 202 and 6.

A few more inputs show the same patterns. Each chip in the last column is one token:

Real splits from the o200k_base tokenizer
TextTokensPieces
The quick brown fox jumps over the lazy dog.10The ␣quick ␣brown ␣fox ␣jumps ␣over ␣the ␣lazy ␣dog .
Tokenization is weird.5Token ization ␣is ␣weird .
unbelievably3un bel ievably
ChatGPT2Chat GPT
2026-10-086202 6 - 10 - 08
12345673123 456 7

␣ marks a space that belongs to the token. Counts from gpt-tokenizer 4.0, o200k_base (GPT-4o to GPT-5.x).

Three things stand out. A token isn’t a word: “Tokenization” is two tokens, and a date is six. The space matters: “unbelievably” on its own is three tokens, but “␣unbelievably” after a space is a single token, because that’s how it usually appears in text. And numbers are split into groups of up to three digits, which is one reason models can stumble on long arithmetic.

How many tokens is a word?

OpenAI’s own rule of thumb is that one token is about four characters or three quarters of a word of English. Our measurement agrees: 1,304 words of ordinary English prose came to 1,628 tokens with o200k_base, which is 0.80 words or 4.8 characters per token. Useful conversions for English:

  • 100 tokens ≈ 75 words ≈ a short paragraph
  • 1,000 tokens ≈ 750 words ≈ a short blog post
  • 100,000 tokens ≈ 75,000 words ≈ a full novel

These are averages for prose. Code, JSON, tables and logs contain lots of punctuation and short symbols, so they use more tokens per word: the same tokenizer needed about 2.1 tokens per whitespace-separated word on a TypeScript file and 3 on a JSON file.

Why other languages use more tokens

Tokenizer vocabularies are built from training text, and most of that text is English. Languages that appear less often, or use other scripts, get fewer whole-word tokens, so the same meaning costs more. Newer tokenizers have narrowed the gap a lot:

Short sentences in four languages
Texto200k_base (GPT-4o to GPT-5.x)cl100k_base (GPT-4, GPT-3.5)
Das ist ein Beispiel. (German)55
यह एक उदाहरण है। (Hindi)520
这是一个例子。 (Chinese)56
مرحبا بالعالم (Arabic)410

The German, Hindi and Chinese sentences all mean “This is an example.”; the Arabic one means “Hello, world”. The Hindi sentence took four times as many tokens on the older tokenizer.

The same four sentences on two generations of OpenAI tokenizer. Newer vocabularies cover more scripts with whole-word tokens.

So if you work in Hindi, Arabic, Urdu or similar languages, the tokenizer generation matters as much as the price per token. Always measure with your own text.

Different models, different counts

Every model family has its own tokenizer, so the same text gives different counts on GPT, Claude, Gemini, Llama or Qwen. A price per million tokens is only comparable once you know how many tokens your text becomes on each model.

  • OpenAI publishes the tokenizers for its models up to GPT-5.x in tiktoken, so those counts can be exact offline.
  • Open-weight models (Llama, Qwen, Mistral, DeepSeek) ship their tokenizer with the weights, so counts can be exact too.
  • Anthropic doesn’t publish Claude’s tokenizer, so any offline number for Claude is an estimate. Its free count_tokens endpoint gives the official count, which Anthropic says can differ slightly from a real request.
  • Google offers a free countTokens endpoint for Gemini, and its Python SDK (google-genai) can also count text locally for supported Gemini models with its LocalTokenizer.

Counts can shift even within one family: Anthropic says Claude Opus 4.7 and later use a newer tokenizer that turns the same text into about 30% more tokens than earlier Claude models. That’s why our token counter doesn’t guess for Claude and Gemini: it shows OpenAI’s count marked for reference only, and fetches the official count when you add your own key.

Free toolAI token counterExact counts for models with public tokenizers, official Claude and Gemini counts with your key, and a view that shows every token.

Why tokens matter: cost and context

Tokens decide two things you care about:

  1. Cost. APIs charge per million input tokens and per million output tokens, and output is usually several times more expensive. The LLM cost calculator turns token counts into a monthly bill.
  2. Context. A model’s context window is the most tokens it can handle in one request: your prompt, the conversation so far, any documents, and the reply. The context window checker tells you whether a document fits.

In a chat or agent, the whole conversation is sent again on every turn. A 20-turn conversation therefore pays for the first message 20 times, which is why long sessions get expensive and why prompt caching exists.

How to use fewer tokens

  • Send only what the model needs: trim logs, strip boilerplate, and summarise long histories instead of resending them.
  • Minify JSON you send as input; whitespace and indentation are tokens too.
  • Ask for short output when that’s all you need. Output tokens cost the most.
  • Use prompt caching for long, repeated prefixes like system prompts and documents.
  • Pick a newer tokenizer for non-English text when you have the choice.

FAQ

Questions people ask

How many tokens is 1,000 words?

About 1,250 tokens for ordinary English prose with OpenAI’s o200k_base tokenizer (we measured 1.25 tokens per word; OpenAI’s rule of thumb gives about 1,330). Code, JSON and non-English text use more. Paste your text into the token counter for the exact number.

Is a token the same as a word?

No. Short common words are usually one token, but longer or rarer words are split into pieces, punctuation is usually its own token, and the space before a word is normally part of that word’s token.

Do spaces and new lines count as tokens?

Yes, though usually not separately. A single space is merged into the next word’s token, while runs of spaces, tabs and blank lines become tokens of their own. Indented code and pretty-printed JSON use noticeably more tokens than minified text.

Why does Claude count my text differently from GPT?

Each model family has its own tokenizer with a different vocabulary. Anthropic doesn’t publish Claude’s tokenizer, so official Claude counts come only from its count_tokens endpoint, which is free to call with an API key. Claude models differ among themselves too: Anthropic says Opus 4.7 and later count about 30% more tokens for the same text.

Do images count as tokens?

Yes. Vision models convert images to tokens too, and the count depends on the image size and the provider’s rules rather than on any text. Providers document their own formulas for images.

Try it

Tools from this guide

Keep reading