AI glossary · Tokens and cost
What is a tokenizer in AI?
Also called: tokeniser, BPE tokenizer
Definition
A tokenizer is the part of a language model that splits text into tokens from a fixed vocabulary, converts them to ID numbers for the model, and turns the model’s output IDs back into text.
Explained
How it works
A tokenizer has two parts: a vocabulary, where every known piece has an ID, and rules for splitting text into those pieces. Most LLM tokenizers use byte-pair encoding (BPE), which builds the vocabulary by repeatedly merging the most frequent pairs of bytes in training text. Because it works on bytes, it can encode any text, even characters it never saw, and decoding gives back exactly the original.
The tokenizer is fixed when the model is trained, and the model only understands its own IDs. OpenAI publishes its tokenizers in tiktoken: o200k_base for GPT-4o and the GPT-5 models, cl100k_base for GPT-4 and GPT-3.5 Turbo. Claude has no public tokenizer, so you count with Anthropic’s free count_tokens endpoint, and Gemini has countTokens.
Vocabulary size matters: the names give it away, with about 200,000 entries in o200k_base and about 100,000 in cl100k_base. A bigger vocabulary holds more whole words, especially outside English, so the same text needs fewer tokens.
Example
The same text on two OpenAI tokenizers
We ran three short texts through both tokenizers. The English sentence splits into the same 7 pieces on each, but the IDs differ: “Token” is 4421 in o200k_base and 3404 in cl100k_base. IDs from one tokenizer mean nothing to a model trained on another.
The Russian sentence needs 6 tokens on the newer vocabulary and 8 on the older one. For Bengali the gap is 11 against 29: cl100k_base has so few Bengali pieces that it falls back to raw bytes, splitting single characters. To count in code, see how to count tokens in Python and JavaScript.
| Text | o200k_base (GPT-4o, GPT-5) | cl100k_base (GPT-4, GPT-3.5) |
|---|---|---|
| Tokenizers turn text into numbers. (English) | 7 | 7 |
| Привет, как дела? (Russian) | 6 | 8 |
| নমস্কার, আপনি কেমন আছেন? (Bengali) | 11 | 29 |
The Russian and Bengali sentences both mean “Hello, how are you?”. Measured with gpt-tokenizer 4.0 on 2026-10-11.
Cost and quality
Why it matters
A price per token only compares fairly between models that share a tokenizer. A model with a lower rate but a hungrier tokenizer can cost more for the same text, and the gap is widest for code and languages other than English.
Tokenizers also change between model generations. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same input than earlier Claude models, so a budget measured on an old model can be badly off. Count with the tokenizer of the model you will actually call.
Don’t mix up
Common confusions
- Tokenizer vs embedding model
- A tokenizer only maps text to ID numbers; it knows nothing about meaning. An embedding model turns text into vectors that capture meaning, and it runs its own tokenizer first.
- Tokenizer vs token counter
- A token counter runs a tokenizer and reports the total. It is exact only when it uses the model’s own tokenizer. Claude’s tokenizer isn’t public, so a tiktoken count for Claude is only a rough guess; Anthropic’s counting API is much closer, though Anthropic itself calls it an estimate that can differ slightly from the billed count.
Go deeper
Try it and read more
- Free toolAI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- Free toolTokens to Words ConverterConvert between tokens, words and pages.
- Guide · 13 min readHow to count tokens in Python and JavaScript (OpenAI, Claude, Gemini, Qwen)Count LLM tokens in Python and JavaScript: tiktoken, gpt-tokenizer, Hugging Face tokenizers, and the official Claude, Gemini and OpenAI count endpoints.
- Guide · 7 min readWhat is a token in AI? A plain-English guide with real examplesWhat LLM tokens are, how text is split into them, why the same text costs different amounts on different models, and how to count tokens exactly.
Related
Related terms
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- Context windowA context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.
- EmbeddingsEmbeddings are lists of numbers (vectors) that an embedding model produces from text, images or other data, placed so that inputs with similar meaning end up close together and can be compared mathematically.
- Open weightsOpen weights means a model’s trained parameters are published for anyone to download and run on their own hardware, under a licence that sets what they may do with them.