Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Guide · Tokens & costs

How to count tokens in Python and JavaScript (OpenAI, Claude, Gemini, Qwen)

To count tokens, run the model’s own tokenizer: tiktoken in Python or gpt-tokenizer in JavaScript for OpenAI models, and the Hugging Face tokenizer for open-weight models like Qwen. Claude has no public tokenizer and Google’s local one doesn’t cover the newest Gemini models, so call their counting endpoints. Count the whole request, not just the text, because chat formatting adds tokens.

By Tahir NazirUpdated 13 min read

On this page
  1. Which token counting method should you use?
  2. How to count tokens in Python with tiktoken
  3. How to count tokens in JavaScript and TypeScript
  4. How to count tokens for Qwen, Llama and other open models
  5. How to count tokens for Claude and Gemini
  6. How to count tokens in a chat conversation
  7. How to count tokens fast
  8. Common mistakes that give the wrong count
  9. Questions people ask

Which token counting method should you use?

Use the model’s own tokenizer when it’s public, and the provider’s counting API when it isn’t. A token count is only exact when it comes from the same tokenizer the model uses, so the right method depends on the model family, not on your programming language:

Pick the lane by model family. Only the first lane runs offline; the second needs an API key; the third is a rough guess you should replace with the response’s usage figures.

If you only need a quick number for a prompt you have open, paste it into the token counter: it runs OpenAI’s tokenizers and several open-weight ones in your browser, and fetches official Claude and Gemini counts with your own key. The rest of this guide is for doing it in code. If tokens themselves are new to you, start with what a token is.

How to count tokens in Python with tiktoken

Install tiktoken, load the encoding for your model, encode the text and take the length of the list. That’s the whole job for plain text on OpenAI models:

count.py
# pip install tiktoken
import tiktoken

text = "Count tokens before you send the request: it saves money."

enc = tiktoken.encoding_for_model("gpt-4o")
print(enc.name, len(enc.encode(text)))

# Load an encoding by name when you know it, or when the model isn't mapped
enc = tiktoken.get_encoding("cl100k_base")
print(enc.name, len(enc.encode(text)))

# Output:
# o200k_base 12
# cl100k_base 12

encoding_for_model looks the name up in a table inside tiktoken and raises a KeyError for any name it doesn’t know. These are the mappings in tiktoken 0.14.0’s source:

Which encoding each OpenAI model uses (tiktoken 0.14.0)
ModelsEncoding
GPT-5 (every name starting gpt-5), GPT-4.5, GPT-4.1, GPT-4o, o1, o3, o4-minio200k_base
gpt-osso200k_harmony
GPT-4, GPT-4 Turbo, GPT-3.5 Turbo, text-embedding-3-small, text-embedding-3-large, text-embedding-ada-002cl100k_base
text-davinci-002, text-davinci-003, Codex modelsp50k_base
GPT-3 base models (davinci, curie, babbage, ada)r50k_base
Anything else, for example gpt-6-astra or any Claude nameKeyError

From MODEL_TO_ENCODING and MODEL_PREFIX_TO_ENCODING in tiktoken/model.py. We ran encoding_for_model on each example name to confirm.

The first call to an encoding downloads its vocabulary file (a few megabytes) from OpenAI’s public storage and caches it in a data-gym-cache folder in your temp directory. For containers, CI or serverless functions without internet access, set TIKTOKEN_CACHE_DIR to a folder you ship with the app and warm it at build time.

How to count tokens in JavaScript and TypeScript

In Node, Deno, Bun, edge runtimes or the browser, use gpt-tokenizer. It’s written in TypeScript, needs no WebAssembly, defaults to o200k_base and gives the same token IDs as tiktoken:

count.ts
// npm install gpt-tokenizer
import { countTokens, isWithinTokenLimit } from "gpt-tokenizer"; // o200k_base
import { countTokens as countCl100k } from "gpt-tokenizer/encoding/cl100k_base";

const text = "Count tokens before you send the request: it saves money.";

console.log(countTokens(text));             // 12
console.log(countCl100k(text));             // 12
console.log(isWithinTokenLimit(text, 10));  // false
console.log(isWithinTokenLimit(text, 100)); // 12

isWithinTokenLimit returns false as soon as the limit is passed, or the count if the text fits, so it’s the cheap way to guard a context limit. js-tiktoken, a community pure-JavaScript port of tiktoken, works too:

count-js-tiktoken.ts
// npm install js-tiktoken
import { getEncoding } from "js-tiktoken";

const enc = getEncoding("o200k_base"); // build once, then reuse
console.log(enc.encode("Count tokens before you send the request: it saves money.").length); // 12

How to count tokens for Qwen, Llama and other open models

Open-weight models ship their tokenizer with the weights, so you can count exactly with Hugging Face’s tokenizers library. Load it straight from the model’s repo; only the tokenizer file is downloaded, not the model:

count_qwen.py
# pip install tokenizers
from tokenizers import Tokenizer

tok = Tokenizer.from_pretrained("Qwen/Qwen3-8B")  # fetches tokenizer.json once
enc = tok.encode("Count tokens before you send the request: it saves money.", add_special_tokens=False)
print(len(enc.ids))
print(enc.tokens[:4])

# Output:
# 12
# ['Count', 'Ġtokens', 'Ġbefore', 'Ġyou']

The Ġ is how byte-level tokenizers show a leading space. add_special_tokens=False counts only your text; some tokenizers otherwise add special tokens such as a beginning-of-sequence marker. If you already use transformers, AutoTokenizer.from_pretrained(...) gives the same counts and adds chat templates, shown in the chat section below.

The Qwen, DeepSeek and Mistral repos we tried download without an account. Meta’s Llama and Google’s Gemma repos are gated: accept the licence on the model page, then log in with hf auth login or set HF_TOKEN. Use the repo for the exact model you call, because tokenizers change between generations.

How to count tokens for Claude and Gemini

Call the provider’s counting endpoint with the same payload you’d send to generate a reply. Anthropic doesn’t publish a tokenizer for current Claude models. Google’s google-genai SDK has an offline LocalTokenizer, but in version 2.29.0 it accepts only some Gemini models (the newest is Gemini 3.5 Flash) and raises “not supported” for gemini-3.8-flash, so for the newest models the endpoint is the only official count. Anthropic’s endpoint is free but rate-limited per usage tier, and Google’s Firebase AI Logic documentation says there’s no charge for calling countTokens. Both need an API key, so keep these calls on a server or in a script, never in public front-end code.

Claude: messages.count_tokens

claude_count.py
# pip install anthropic   (reads ANTHROPIC_API_KEY)
import anthropic

client = anthropic.Anthropic()

response = client.messages.count_tokens(
    model="claude-opus-5-5",
    system="You are a scientist",
    messages=[{"role": "user", "content": "Hello, Claude"}],
)

print(response.json())
claude-count.ts
// npm install @anthropic-ai/sdk   (reads ANTHROPIC_API_KEY)
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

const response = await client.messages.countTokens({
  model: "claude-opus-5-5",
  system: "You are a scientist",
  messages: [
    {
      role: "user",
      content: "Hello, Claude"
    }
  ]
});

console.log(response);

Anthropic’s documentation shows { "input_tokens": 14 } for this request. It also says the count is an estimate that can differ from the billed input by a small amount, and that Claude Opus 4.7 and later use a newer tokenizer that produces roughly 30% more tokens for the same text. Count against the exact model you’ll call, and don’t reuse counts across Claude generations.

Gemini: models.countTokens

gemini_count.py
# pip install google-genai   (reads GEMINI_API_KEY)
from google import genai

client = genai.Client()
prompt = "The quick brown fox jumps over the lazy dog."

total_tokens = client.models.count_tokens(
    model="gemini-3.8-flash",
    contents=prompt
)
print("total_tokens:", total_tokens.total_tokens)
gemini-count.ts
// npm install @google/genai   (reads GEMINI_API_KEY)
import { GoogleGenAI } from "@google/genai";

const client = new GoogleGenAI({});
const prompt = "The quick brown fox jumps over the lazy dog.";

const countResponse = await client.models.countTokens({
  model: "gemini-3.8-flash",
  contents: prompt,
});
console.log(countResponse.totalTokens);

OpenAI has a counting endpoint too, responses.input_tokens.count. Use it when a local count isn’t enough: requests with tools, images or files, and models tiktoken doesn’t map yet. OpenAI says it returns the exact count the model will receive, formatting tokens included:

openai_count.py
# pip install openai   (reads OPENAI_API_KEY)
from openai import OpenAI

client = OpenAI()

response = client.responses.input_tokens.count(
    model="gpt-6-astra", input="Tell me a joke."
)
print(response.input_tokens)
Free toolAI token counterCompare counts across GPT, Claude, Gemini, Qwen and DeepSeek in one view. Add your own key for official Claude and Gemini counts; it goes straight from your browser to the provider.

How to count tokens in a chat conversation

Count the formatted request, not just the message text. Every chat model wraps each message in role markers and separators, and those are tokens you pay for. Take this two-message conversation:

  • System: “You are a helpful, pattern-following assistant that translates corporate jargon into plain English.”
  • User: “This late pivot means we don’t have time to boil the ocean for the client deliverable.”

Its message text and role names come to 37 tokens with o200k_base. For OpenAI’s Chat Completions format, OpenAI’s cookbook gives the rule: 3 tokens per message, 1 more for a name field, and 3 to start the assistant’s reply. The cookbook checked that rule against the API for GPT-3.5 Turbo, GPT-4, GPT-4o and GPT-4o mini:

chat_tokens.py
import tiktoken

def num_tokens_from_messages(messages, encoding_name="o200k_base"):
    """OpenAI cookbook rule for Chat Completions: 3 tokens per message,
    1 more per name field, and 3 to prime the assistant's reply."""
    enc = tiktoken.get_encoding(encoding_name)
    num_tokens = 0
    for message in messages:
        num_tokens += 3
        for key, value in message.items():
            num_tokens += len(enc.encode(value))
            if key == "name":
                num_tokens += 1
    return num_tokens + 3

messages = [
    {"role": "system", "content": "You are a helpful, pattern-following assistant that translates corporate jargon into plain English."},
    {"role": "user", "content": "This late pivot means we don't have time to boil the ocean for the client deliverable."},
]
print(num_tokens_from_messages(messages))  # 46

In JavaScript, encodeChat from gpt-tokenizer applies the same formatting and agrees:

chat-tokens.ts
import { encodeChat } from "gpt-tokenizer";

const chat = [
  { role: "system", content: "You are a helpful, pattern-following assistant that translates corporate jargon into plain English." },
  { role: "user", content: "This late pivot means we don't have time to boil the ocean for the client deliverable." },
];
console.log(encodeChat(chat, "gpt-4o").length); // 46

Open-weight models format chats with a template stored in the repo. apply_chat_template renders it, so the count includes every marker. It needs jinja2 installed; without it, transformers raises an ImportError:

chat_tokens_qwen.py
# pip install transformers jinja2
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
messages = [
    {"role": "system", "content": "You are a helpful, pattern-following assistant that translates corporate jargon into plain English."},
    {"role": "user", "content": "This late pivot means we don't have time to boil the ocean for the client deliverable."},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=True, return_dict=False)
print(len(ids))  # 50
The same two messages, counted three ways
MethodTokens
Message text and role names, o200k_base37
GPT-4o request (cookbook rule, and gpt-tokenizer’s encodeChat)46
Qwen3-8B with its chat template50

Measured on 2026-10-08. Qwen3’s tokenizer splits the two message texts into 37 tokens (o200k_base: 35), so its template, role names included, adds 13.

Two limits apply. The cookbook rule covers Chat Completions on those older models, and OpenAI notes that tool definitions and model-specific formatting add tokens that are hard to count locally, so use responses.input_tokens.count for anything with tools. For Claude and Gemini, send the whole conversation, system prompt and tools to the counting endpoint, which counts the formatting for you. And remember the conversation is sent again on every turn, so a long chat’s cost grows with each reply: the LLM cost calculator shows what that does to a monthly bill, and the guide to estimating LLM API costs walks through the sums.

How to count tokens fast

Local counting is fast enough for almost any app if you load the tokenizer once and reuse it. We counted the same 1,000,000-character Markdown file with each library:

Counting 1,000,000 characters on a laptop
LibraryTokensTime
tiktoken 0.14.0, o200k_base (Python)259,4010.07 to 0.11 s
gpt-tokenizer 4.0.0, o200k_base (Node)259,4010.13 to 0.35 s
js-tiktoken 1.0.21, o200k_base (Node)259,4010.6 to 1.7 s, plus 0.7 to 2 s to build the encoder
tokenizers 0.23.2, Qwen3-8B, one string (Python)266,3931.0 to 1.4 s
tokenizers 0.23.2, Qwen3-8B, encode_batch over 12,497 lines264,9690.26 to 0.41 s

Each figure is the median of five runs on an Intel Core i7-1185G7, measured 2026-10-08; ranges span repeated sessions. Absolute times depend on your machine; the order is what matters.

  • Load once. Create the encoder or tokenizer at start-up, not per request. js-tiktoken spent longer building its encoder than counting a million characters.
  • Batch for Hugging Face. encode_batch was more than three times faster than one huge string. With tiktoken, a single encode call on the whole text was the fastest option we measured.
  • Stop early. For a yes-or-no limit check, gpt-tokenizer’s isWithinTokenLimit stops as soon as the limit is passed.
  • Cache counts. A document’s count doesn’t change; store it next to the document instead of recounting on every request.
  • Mind the browser bundle. gpt-tokenizer was the fastest JavaScript option here and needs no WebAssembly. Load it with a dynamic import() so it only downloads when needed.

Common mistakes that give the wrong count

  1. Using one tokenizer for every model. tiktoken is exact for OpenAI only. On this 1,000,000-character file, Qwen3 needed 2.7% more tokens than o200k_base; on non-English text the gap is much wider. A tiktoken number for Claude or Gemini is an estimate.
  2. Counting text, not the request. System prompts, chat formatting, tool schemas and images all add input tokens. Count the payload you actually send.
  3. Choking on special tokens. If user text contains a string like <|endoftext|>, tiktoken’s encode raises a ValueError and gpt-tokenizer throws “Disallowed special token found”. To count it as ordinary text, use enc.encode(text, disallowed_special=()) or enc.encode_ordinary(text) in Python (9 tokens for Ignore <|endoftext|> this), or encode(text, { disallowedSpecial: new Set() }) in gpt-tokenizer.
  4. Assuming the model name is mapped. New model names raise KeyError in tiktoken or “Unknown model” in js-tiktoken until the library is updated. Pin the encoding by name, or use the provider’s counting endpoint.
  5. Forgetting output tokens. Input counts don’t include the reply. OpenAI notes that reported output can be higher than the visible text even when no reasoning tokens are reported, because some models generate formatting tokens you don’t see, so leave headroom in max_output_tokens.
  6. Reusing old counts after a model change. Anthropic’s newer tokenizer is the clearest example: the same prompt counts about 30% higher on Claude Opus 4.7 and later. Recount when you switch models, then check the result still fits with the context window checker.

FAQ

Questions people ask

How do I count tokens without tiktoken?

In JavaScript, use gpt-tokenizer or js-tiktoken, which give the same counts as tiktoken for OpenAI encodings. For open-weight models, use Hugging Face’s tokenizers or transformers. For Claude and Gemini, call count_tokens or countTokens. Without any library, about four characters per token is the usual English rule of thumb, but it’s only an estimate.

Is tiktoken accurate for Claude, Gemini or Llama?

No. tiktoken implements OpenAI’s encodings only. Claude, Gemini, Llama and Qwen each have their own tokenizer, so a tiktoken count for them is a rough estimate that can be noticeably off, especially for code and non-English text. Use the Hugging Face tokenizer for open-weight models and the official counting endpoints for Claude and Gemini.

Does tiktoken work offline?

Yes, after the first run. Each encoding’s vocabulary file is downloaded on first use and cached in a data-gym-cache folder in your temp directory. For offline machines, containers or serverless functions, set TIKTOKEN_CACHE_DIR to a folder you ship with the app and populate it during your build.

Is Anthropic’s count_tokens endpoint free?

Yes. Anthropic says token counting is free to use, with requests-per-minute limits that depend on your usage tier and are separate from your message limits. It needs an API key. Anthropic also notes the result is an estimate that can differ slightly from the billed input, and it doesn’t accept requests that use server tools such as web search.

How many tokens does a chat message add?

For OpenAI Chat Completions on models like GPT-4o, OpenAI’s cookbook counts 3 tokens per message, 1 more for a name, and 3 to start the reply. Our two-message example went from 37 tokens of text to 46. Open-weight chat templates differ: the same messages were 50 tokens on Qwen3-8B.

Why does my count differ from the usage the API reports?

Usually because you counted only the text. The API also counts message formatting, the system prompt, tool definitions and images. A wrong encoding, or a model the library doesn’t know, causes the rest. For Claude, Anthropic says the counting endpoint itself is an estimate, so small differences are expected.

Try it

Tools from this guide

Keep reading