Guide · Tokens & costs
How to count tokens in Python and JavaScript (OpenAI, Claude, Gemini, Qwen)
To count tokens, run the model’s own tokenizer: tiktoken in Python or gpt-tokenizer in JavaScript for OpenAI models, and the Hugging Face tokenizer for open-weight models like Qwen. Claude has no public tokenizer and Google’s local one doesn’t cover the newest Gemini models, so call their counting endpoints. Count the whole request, not just the text, because chat formatting adds tokens.
By Tahir NazirUpdated 13 min read
On this page
- Which token counting method should you use?
- How to count tokens in Python with tiktoken
- How to count tokens in JavaScript and TypeScript
- How to count tokens for Qwen, Llama and other open models
- How to count tokens for Claude and Gemini
- How to count tokens in a chat conversation
- How to count tokens fast
- Common mistakes that give the wrong count
- Questions people ask
Which token counting method should you use?
Use the model’s own tokenizer when it’s public, and the provider’s counting API when it isn’t. A token count is only exact when it comes from the same tokenizer the model uses, so the right method depends on the model family, not on your programming language:
- OpenAI GPT-5, GPT-4.1, GPT-4o, o-serieso200k_basetiktoken · gpt-tokenizer · js-tiktoken
- OpenAI GPT-4, GPT-3.5cl100k_basetiktoken · gpt-tokenizer · js-tiktoken
- OpenAI gpt-osso200k_harmonytiktoken · gpt-tokenizer
- Open-weight: Qwen, Llama, DeepSeek, Mistral, Gemmatokenizer.json from the model’s repotokenizers · transformers
- Anthropic Claudemessages.count_tokensanthropic · @anthropic-ai/sdk
- Google Geminimodels.countTokensgoogle-genai · @google/genai
- OpenAI requests with tools, images or files, and unmapped modelsresponses.input_tokens.countopenai
- Any model with no public tokenizer and no count endpoint≈ 4 characters per English tokenreplace with usage from the response
If you only need a quick number for a prompt you have open, paste it into the token counter: it runs OpenAI’s tokenizers and several open-weight ones in your browser, and fetches official Claude and Gemini counts with your own key. The rest of this guide is for doing it in code. If tokens themselves are new to you, start with what a token is.
How to count tokens in Python with tiktoken
Install tiktoken, load the encoding for your model, encode the text and take the length of the list. That’s the whole job for plain text on OpenAI models:
# pip install tiktoken
import tiktoken
text = "Count tokens before you send the request: it saves money."
enc = tiktoken.encoding_for_model("gpt-4o")
print(enc.name, len(enc.encode(text)))
# Load an encoding by name when you know it, or when the model isn't mapped
enc = tiktoken.get_encoding("cl100k_base")
print(enc.name, len(enc.encode(text)))
# Output:
# o200k_base 12
# cl100k_base 12encoding_for_model looks the name up in a table inside tiktoken and raises a KeyError for any name it doesn’t know. These are the mappings in tiktoken 0.14.0’s source:
| Models | Encoding |
|---|---|
GPT-5 (every name starting gpt-5), GPT-4.5, GPT-4.1, GPT-4o, o1, o3, o4-mini | o200k_base |
| gpt-oss | o200k_harmony |
GPT-4, GPT-4 Turbo, GPT-3.5 Turbo, text-embedding-3-small, text-embedding-3-large, text-embedding-ada-002 | cl100k_base |
text-davinci-002, text-davinci-003, Codex models | p50k_base |
GPT-3 base models (davinci, curie, babbage, ada) | r50k_base |
Anything else, for example gpt-6-astra or any Claude name | KeyError |
From MODEL_TO_ENCODING and MODEL_PREFIX_TO_ENCODING in tiktoken/model.py. We ran encoding_for_model on each example name to confirm.
The first call to an encoding downloads its vocabulary file (a few megabytes) from OpenAI’s public storage and caches it in a data-gym-cache folder in your temp directory. For containers, CI or serverless functions without internet access, set TIKTOKEN_CACHE_DIR to a folder you ship with the app and warm it at build time.
How to count tokens in JavaScript and TypeScript
In Node, Deno, Bun, edge runtimes or the browser, use gpt-tokenizer. It’s written in TypeScript, needs no WebAssembly, defaults to o200k_base and gives the same token IDs as tiktoken:
// npm install gpt-tokenizer
import { countTokens, isWithinTokenLimit } from "gpt-tokenizer"; // o200k_base
import { countTokens as countCl100k } from "gpt-tokenizer/encoding/cl100k_base";
const text = "Count tokens before you send the request: it saves money.";
console.log(countTokens(text)); // 12
console.log(countCl100k(text)); // 12
console.log(isWithinTokenLimit(text, 10)); // false
console.log(isWithinTokenLimit(text, 100)); // 12isWithinTokenLimit returns false as soon as the limit is passed, or the count if the text fits, so it’s the cheap way to guard a context limit. js-tiktoken, a community pure-JavaScript port of tiktoken, works too:
// npm install js-tiktoken
import { getEncoding } from "js-tiktoken";
const enc = getEncoding("o200k_base"); // build once, then reuse
console.log(enc.encode("Count tokens before you send the request: it saves money.").length); // 12How to count tokens for Qwen, Llama and other open models
Open-weight models ship their tokenizer with the weights, so you can count exactly with Hugging Face’s tokenizers library. Load it straight from the model’s repo; only the tokenizer file is downloaded, not the model:
# pip install tokenizers
from tokenizers import Tokenizer
tok = Tokenizer.from_pretrained("Qwen/Qwen3-8B") # fetches tokenizer.json once
enc = tok.encode("Count tokens before you send the request: it saves money.", add_special_tokens=False)
print(len(enc.ids))
print(enc.tokens[:4])
# Output:
# 12
# ['Count', 'Ġtokens', 'Ġbefore', 'Ġyou']The Ġ is how byte-level tokenizers show a leading space. add_special_tokens=False counts only your text; some tokenizers otherwise add special tokens such as a beginning-of-sequence marker. If you already use transformers, AutoTokenizer.from_pretrained(...) gives the same counts and adds chat templates, shown in the chat section below.
The Qwen, DeepSeek and Mistral repos we tried download without an account. Meta’s Llama and Google’s Gemma repos are gated: accept the licence on the model page, then log in with hf auth login or set HF_TOKEN. Use the repo for the exact model you call, because tokenizers change between generations.
How to count tokens for Claude and Gemini
Call the provider’s counting endpoint with the same payload you’d send to generate a reply. Anthropic doesn’t publish a tokenizer for current Claude models. Google’s google-genai SDK has an offline LocalTokenizer, but in version 2.29.0 it accepts only some Gemini models (the newest is Gemini 3.5 Flash) and raises “not supported” for gemini-3.8-flash, so for the newest models the endpoint is the only official count. Anthropic’s endpoint is free but rate-limited per usage tier, and Google’s Firebase AI Logic documentation says there’s no charge for calling countTokens. Both need an API key, so keep these calls on a server or in a script, never in public front-end code.
Claude: messages.count_tokens
# pip install anthropic (reads ANTHROPIC_API_KEY)
import anthropic
client = anthropic.Anthropic()
response = client.messages.count_tokens(
model="claude-opus-5-5",
system="You are a scientist",
messages=[{"role": "user", "content": "Hello, Claude"}],
)
print(response.json())// npm install @anthropic-ai/sdk (reads ANTHROPIC_API_KEY)
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const response = await client.messages.countTokens({
model: "claude-opus-5-5",
system: "You are a scientist",
messages: [
{
role: "user",
content: "Hello, Claude"
}
]
});
console.log(response);Anthropic’s documentation shows { "input_tokens": 14 } for this request. It also says the count is an estimate that can differ from the billed input by a small amount, and that Claude Opus 4.7 and later use a newer tokenizer that produces roughly 30% more tokens for the same text. Count against the exact model you’ll call, and don’t reuse counts across Claude generations.
Gemini: models.countTokens
# pip install google-genai (reads GEMINI_API_KEY)
from google import genai
client = genai.Client()
prompt = "The quick brown fox jumps over the lazy dog."
total_tokens = client.models.count_tokens(
model="gemini-3.8-flash",
contents=prompt
)
print("total_tokens:", total_tokens.total_tokens)// npm install @google/genai (reads GEMINI_API_KEY)
import { GoogleGenAI } from "@google/genai";
const client = new GoogleGenAI({});
const prompt = "The quick brown fox jumps over the lazy dog.";
const countResponse = await client.models.countTokens({
model: "gemini-3.8-flash",
contents: prompt,
});
console.log(countResponse.totalTokens);OpenAI has a counting endpoint too, responses.input_tokens.count. Use it when a local count isn’t enough: requests with tools, images or files, and models tiktoken doesn’t map yet. OpenAI says it returns the exact count the model will receive, formatting tokens included:
# pip install openai (reads OPENAI_API_KEY)
from openai import OpenAI
client = OpenAI()
response = client.responses.input_tokens.count(
model="gpt-6-astra", input="Tell me a joke."
)
print(response.input_tokens)How to count tokens in a chat conversation
Count the formatted request, not just the message text. Every chat model wraps each message in role markers and separators, and those are tokens you pay for. Take this two-message conversation:
- System: “You are a helpful, pattern-following assistant that translates corporate jargon into plain English.”
- User: “This late pivot means we don’t have time to boil the ocean for the client deliverable.”
Its message text and role names come to 37 tokens with o200k_base. For OpenAI’s Chat Completions format, OpenAI’s cookbook gives the rule: 3 tokens per message, 1 more for a name field, and 3 to start the assistant’s reply. The cookbook checked that rule against the API for GPT-3.5 Turbo, GPT-4, GPT-4o and GPT-4o mini:
import tiktoken
def num_tokens_from_messages(messages, encoding_name="o200k_base"):
"""OpenAI cookbook rule for Chat Completions: 3 tokens per message,
1 more per name field, and 3 to prime the assistant's reply."""
enc = tiktoken.get_encoding(encoding_name)
num_tokens = 0
for message in messages:
num_tokens += 3
for key, value in message.items():
num_tokens += len(enc.encode(value))
if key == "name":
num_tokens += 1
return num_tokens + 3
messages = [
{"role": "system", "content": "You are a helpful, pattern-following assistant that translates corporate jargon into plain English."},
{"role": "user", "content": "This late pivot means we don't have time to boil the ocean for the client deliverable."},
]
print(num_tokens_from_messages(messages)) # 46In JavaScript, encodeChat from gpt-tokenizer applies the same formatting and agrees:
import { encodeChat } from "gpt-tokenizer";
const chat = [
{ role: "system", content: "You are a helpful, pattern-following assistant that translates corporate jargon into plain English." },
{ role: "user", content: "This late pivot means we don't have time to boil the ocean for the client deliverable." },
];
console.log(encodeChat(chat, "gpt-4o").length); // 46Open-weight models format chats with a template stored in the repo. apply_chat_template renders it, so the count includes every marker. It needs jinja2 installed; without it, transformers raises an ImportError:
# pip install transformers jinja2
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
messages = [
{"role": "system", "content": "You are a helpful, pattern-following assistant that translates corporate jargon into plain English."},
{"role": "user", "content": "This late pivot means we don't have time to boil the ocean for the client deliverable."},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=True, return_dict=False)
print(len(ids)) # 50| Method | Tokens |
|---|---|
Message text and role names, o200k_base | 37 |
GPT-4o request (cookbook rule, and gpt-tokenizer’s encodeChat) | 46 |
| Qwen3-8B with its chat template | 50 |
Measured on 2026-10-08. Qwen3’s tokenizer splits the two message texts into 37 tokens (o200k_base: 35), so its template, role names included, adds 13.
Two limits apply. The cookbook rule covers Chat Completions on those older models, and OpenAI notes that tool definitions and model-specific formatting add tokens that are hard to count locally, so use responses.input_tokens.count for anything with tools. For Claude and Gemini, send the whole conversation, system prompt and tools to the counting endpoint, which counts the formatting for you. And remember the conversation is sent again on every turn, so a long chat’s cost grows with each reply: the LLM cost calculator shows what that does to a monthly bill, and the guide to estimating LLM API costs walks through the sums.
How to count tokens fast
Local counting is fast enough for almost any app if you load the tokenizer once and reuse it. We counted the same 1,000,000-character Markdown file with each library:
| Library | Tokens | Time |
|---|---|---|
tiktoken 0.14.0, o200k_base (Python) | 259,401 | 0.07 to 0.11 s |
gpt-tokenizer 4.0.0, o200k_base (Node) | 259,401 | 0.13 to 0.35 s |
js-tiktoken 1.0.21, o200k_base (Node) | 259,401 | 0.6 to 1.7 s, plus 0.7 to 2 s to build the encoder |
| tokenizers 0.23.2, Qwen3-8B, one string (Python) | 266,393 | 1.0 to 1.4 s |
tokenizers 0.23.2, Qwen3-8B, encode_batch over 12,497 lines | 264,969 | 0.26 to 0.41 s |
Each figure is the median of five runs on an Intel Core i7-1185G7, measured 2026-10-08; ranges span repeated sessions. Absolute times depend on your machine; the order is what matters.
- Load once. Create the encoder or tokenizer at start-up, not per request. js-tiktoken spent longer building its encoder than counting a million characters.
- Batch for Hugging Face.
encode_batchwas more than three times faster than one huge string. With tiktoken, a singleencodecall on the whole text was the fastest option we measured. - Stop early. For a yes-or-no limit check, gpt-tokenizer’s
isWithinTokenLimitstops as soon as the limit is passed. - Cache counts. A document’s count doesn’t change; store it next to the document instead of recounting on every request.
- Mind the browser bundle. gpt-tokenizer was the fastest JavaScript option here and needs no WebAssembly. Load it with a dynamic
import()so it only downloads when needed.
Common mistakes that give the wrong count
- Using one tokenizer for every model. tiktoken is exact for OpenAI only. On this 1,000,000-character file, Qwen3 needed 2.7% more tokens than
o200k_base; on non-English text the gap is much wider. A tiktoken number for Claude or Gemini is an estimate. - Counting text, not the request. System prompts, chat formatting, tool schemas and images all add input tokens. Count the payload you actually send.
- Choking on special tokens. If user text contains a string like
<|endoftext|>, tiktoken’sencoderaises aValueErrorand gpt-tokenizer throws “Disallowed special token found”. To count it as ordinary text, useenc.encode(text, disallowed_special=())orenc.encode_ordinary(text)in Python (9 tokens forIgnore <|endoftext|> this), orencode(text, { disallowedSpecial: new Set() })in gpt-tokenizer. - Assuming the model name is mapped. New model names raise
KeyErrorin tiktoken or “Unknown model” in js-tiktoken until the library is updated. Pin the encoding by name, or use the provider’s counting endpoint. - Forgetting output tokens. Input counts don’t include the reply. OpenAI notes that reported output can be higher than the visible text even when no reasoning tokens are reported, because some models generate formatting tokens you don’t see, so leave headroom in
max_output_tokens. - Reusing old counts after a model change. Anthropic’s newer tokenizer is the clearest example: the same prompt counts about 30% higher on Claude Opus 4.7 and later. Recount when you switch models, then check the result still fits with the context window checker.
FAQ
Questions people ask
How do I count tokens without tiktoken?
In JavaScript, use gpt-tokenizer or js-tiktoken, which give the same counts as tiktoken for OpenAI encodings. For open-weight models, use Hugging Face’s tokenizers or transformers. For Claude and Gemini, call count_tokens or countTokens. Without any library, about four characters per token is the usual English rule of thumb, but it’s only an estimate.
Is tiktoken accurate for Claude, Gemini or Llama?
No. tiktoken implements OpenAI’s encodings only. Claude, Gemini, Llama and Qwen each have their own tokenizer, so a tiktoken count for them is a rough estimate that can be noticeably off, especially for code and non-English text. Use the Hugging Face tokenizer for open-weight models and the official counting endpoints for Claude and Gemini.
Does tiktoken work offline?
Yes, after the first run. Each encoding’s vocabulary file is downloaded on first use and cached in a data-gym-cache folder in your temp directory. For offline machines, containers or serverless functions, set TIKTOKEN_CACHE_DIR to a folder you ship with the app and populate it during your build.
Is Anthropic’s count_tokens endpoint free?
Yes. Anthropic says token counting is free to use, with requests-per-minute limits that depend on your usage tier and are separate from your message limits. It needs an API key. Anthropic also notes the result is an estimate that can differ slightly from the billed input, and it doesn’t accept requests that use server tools such as web search.
How many tokens does a chat message add?
For OpenAI Chat Completions on models like GPT-4o, OpenAI’s cookbook counts 3 tokens per message, 1 more for a name, and 3 to start the reply. Our two-message example went from 37 tokens of text to 46. Open-weight chat templates differ: the same messages were 50 tokens on Qwen3-8B.
Why does my count differ from the usage the API reports?
Usually because you counted only the text. The API also counts message formatting, the system prompt, tool definitions and images. A wrong encoding, or a model the library doesn’t know, causes the rest. For Claude, Anthropic says the counting endpoint itself is an estimate, so small differences are expected.
Try it
Tools from this guide
Keep reading