AI glossary · Tokens and cost
What does token mean in AI?
Also called: LLM token, AI token
Definition
A token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
Explained
How it works
Before a model sees your prompt, a tokenizer cuts the text into pieces from a fixed vocabulary and swaps each piece for its ID number. The model then predicts one new token at a time, and the tokenizer turns the IDs back into text. The token is the piece; the ID is what the model actually processes.
Common words get a token of their own, usually with the space before them included. Rare words, names and code are built from several pieces, and punctuation is usually separate. Anthropic puts a token at about 3.5 English characters on earlier Claude models, and OpenAI’s tiktoken notes about 4 bytes per token on average.
Tokens are the unit of everything you pay for and every limit you hit: API prices are per million tokens, and a model’s context window and output limit are counted in tokens. Each model family has its own tokenizer, so the same text gives a different count on Claude, GPT and Gemini.
Example
A 7-word sentence, 11 tokens
OpenAI’s o200k_base tokenizer, used by GPT-4o and the GPT-5 models, splits the sentence below into 11 tokens. “Chatbots” becomes two pieces, “don’t” splits into “␣don” and “’t”, and the semicolon and full stop are tokens of their own.
Count your own text with the token counter, which shows the split for several models side by side. For why other languages, numbers and code cost more tokens, read what is a token in AI.
| Result | |
|---|---|
| Pieces | Chat bots ␣don ’t ␣read ␣words ; ␣they ␣read ␣tokens . |
| Token IDs | 14065 91601 1700 1573 1729 6391 26 1023 1729 20290 13 |
| Words | 7 |
| Tokens | 11 |
␣ marks a space that belongs to the token. Measured with gpt-tokenizer 4.0 (o200k_base) on 2026-10-11.
Cost and quality
Why it matters
Every cost and limit is counted in tokens, so a word count is only an estimate. Code, JSON and many languages other than English use more tokens per word than English prose, and output tokens usually cost several times as much as input tokens.
The count also depends on the model. Anthropic says Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text than earlier Claude models, so recount when you switch.
Don’t mix up
Common confusions
- Tokens vs words
- A token is not a word. In the example, 7 words became 11 tokens. In English prose a token averages roughly three quarters of a word, but code, numbers and other languages are denser.
- AI tokens vs access tokens and crypto tokens
- An access token (for example from OAuth) is a credential that proves who you are, and a crypto token is a digital asset. When an AI price says “per million tokens”, it means pieces of text.
Go deeper
Try it and read more
- Free toolAI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- Free toolTokens to Words ConverterConvert between tokens, words and pages.
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Guide · 7 min readWhat is a token in AI? A plain-English guide with real examplesWhat LLM tokens are, how text is split into them, why the same text costs different amounts on different models, and how to count tokens exactly.
- Guide · 13 min readHow to count tokens in Python and JavaScript (OpenAI, Claude, Gemini, Qwen)Count LLM tokens in Python and JavaScript: tiktoken, gpt-tokenizer, Hugging Face tokenizers, and the official Claude, Gemini and OpenAI count endpoints.
Related
Related terms
- TokenizerA tokenizer is the part of a language model that splits text into tokens from a fixed vocabulary, converts them to ID numbers for the model, and turns the model’s output IDs back into text.
- Context windowA context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.
- Max output tokensMax output tokens is the cap on how many tokens a model may generate in one response, set per request up to the model’s own limit, with any reasoning tokens counted inside it.
- Reasoning tokensReasoning tokens are the tokens a reasoning model generates while it works through a problem before writing its answer, billed as output tokens even though the API hides them or returns only a summary.
- Prompt cachingPrompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.