Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Tokens and cost

What is a tokenizer in AI?

Also called: tokeniser, BPE tokenizer

Definition

A tokenizer is the part of a language model that splits text into tokens from a fixed vocabulary, converts them to ID numbers for the model, and turns the model’s output IDs back into text.

Explained

How it works

A tokenizer has two parts: a vocabulary, where every known piece has an ID, and rules for splitting text into those pieces. Most LLM tokenizers use byte-pair encoding (BPE), which builds the vocabulary by repeatedly merging the most frequent pairs of bytes in training text. Because it works on bytes, it can encode any text, even characters it never saw, and decoding gives back exactly the original.

The tokenizer is fixed when the model is trained, and the model only understands its own IDs. OpenAI publishes its tokenizers in tiktoken: o200k_base for GPT-4o and the GPT-5 models, cl100k_base for GPT-4 and GPT-3.5 Turbo. Claude has no public tokenizer, so you count with Anthropic’s free count_tokens endpoint, and Gemini has countTokens.

Vocabulary size matters: the names give it away, with about 200,000 entries in o200k_base and about 100,000 in cl100k_base. A bigger vocabulary holds more whole words, especially outside English, so the same text needs fewer tokens.

Example

The same text on two OpenAI tokenizers

We ran three short texts through both tokenizers. The English sentence splits into the same 7 pieces on each, but the IDs differ: “Token” is 4421 in o200k_base and 3404 in cl100k_base. IDs from one tokenizer mean nothing to a model trained on another.

The Russian sentence needs 6 tokens on the newer vocabulary and 8 on the older one. For Bengali the gap is 11 against 29: cl100k_base has so few Bengali pieces that it falls back to raw bytes, splitting single characters. To count in code, see how to count tokens in Python and JavaScript.

Tokens per text, measured
Texto200k_base (GPT-4o, GPT-5)cl100k_base (GPT-4, GPT-3.5)
Tokenizers turn text into numbers. (English)77
Привет, как дела? (Russian)68
নমস্কার, আপনি কেমন আছেন? (Bengali)1129

The Russian and Bengali sentences both mean “Hello, how are you?”. Measured with gpt-tokenizer 4.0 on 2026-10-11.

Cost and quality

Why it matters

A price per token only compares fairly between models that share a tokenizer. A model with a lower rate but a hungrier tokenizer can cost more for the same text, and the gap is widest for code and languages other than English.

Tokenizers also change between model generations. Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same input than earlier Claude models, so a budget measured on an old model can be badly off. Count with the tokenizer of the model you will actually call.

Don’t mix up

Common confusions

Tokenizer vs embedding model
A tokenizer only maps text to ID numbers; it knows nothing about meaning. An embedding model turns text into vectors that capture meaning, and it runs its own tokenizer first.
Tokenizer vs token counter
A token counter runs a tokenizer and reports the total. It is exact only when it uses the model’s own tokenizer. Claude’s tokenizer isn’t public, so a tiktoken count for Claude is only a rough guess; Anthropic’s counting API is much closer, though Anthropic itself calls it an estimate that can differ slightly from the billed count.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary

Written by Tahir Nazir. Checked .

How this was checked: Token counts and IDs measured with gpt-tokenizer 4.0 (o200k_base and cl100k_base) on 2026-10-11 and re-checked by a unit test, which also checks both vocabulary sizes. BPE properties checked against the tiktoken README, Claude counting and the newer Claude tokenizer against Anthropic’s token counting docs, on 2026-10-11.