Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

Tokens & Costs

Tokens to words converter

How many words is 1,000 tokens? About 866 words of English prose with GPT-5’s tokenizer, but it depends on the language, the kind of text and the model. This free converter uses ratios we measured on real text, and it runs in your browser.

The Universal Declaration of Human Rights in the language you pick

GPT-5, GPT-4.1, GPT-4o and o-series models

A common convention for 12-point text with 1-inch margins

1,000 tokens is about866 words
Pages (500 words)1.7
Characters4,408

Measured: 2,017 OpenAI o200k_base tokens for 1,747 words of English prose. Characters don’t include spaces.

These are averages for this kind of text. For your own text, the token counter gives exact counts per model.

1,000 tokens of English prose, by tokenizer

TokenizerWordsPagesCharacters
OpenAI o200k_baseGPT-5, GPT-4.1, GPT-4o and o-series models8661.74,408
OpenAI cl100k_baseGPT-4 and GPT-3.5 Turbo8671.74,410
DeepSeek V4DeepSeek V4 Pro and V4.1 Flash8731.74,443
Qwen 3.8Qwen 3.8 models8351.74,250
Mistral Small 4Mistral Small 48491.74,320
GeminiGemini models (Google’s countTokens API)8431.74,291
Claude (Opus 4.7 and later)Anthropic’s figure, English only5551.12,825
Claude (before Opus 4.7)Anthropic’s figure, English only7501.53,817

Measured 2026-10-11. “–” means no measurement or published figure for this text.

Check the ratio for your own text

Steps

How to use the tokens to words converter

  1. Type an amount and choose what it is: tokens, words, pages or characters.
  2. Pick the type of text. For prose, pick the language too.
  3. Pick the tokenizer for the model you use. OpenAI’s o200k_base covers GPT-5, GPT-4.1 and GPT-4o.
  4. Read the result, and compare it with the other tokenizers in the table below it.
  5. For an exact count of your own text, open “Check the ratio for your own text” or use the token counter.

Method

How it works

Models read text as tokens: whole words, pieces of words, punctuation and spaces. The popular rule of thumb, from OpenAI, is that a token is about four characters or three-quarters of an English word. Most converters apply that one ratio to everything. It’s in the right range for plain English: with OpenAI’s o200k_base tokenizer we measured 0.87 words per token on formal prose and 0.70 on technical writing. But it can be badly wrong for other languages, for code and data, and for models whose tokenizers split text differently. So this converter uses a measured ratio for each combination.

Why the ratio changes

The tokenizer. Each model family learns its own vocabulary. Larger, newer vocabularies hold more whole words, especially outside English: OpenAI’s older cl100k_base (GPT-4) needed 3.3× as many tokens as o200k_base for the Hindi declaration, while for English the two were almost level (2,016 and 2,017 tokens). Anthropic says its tokenizer for Claude Opus 4.7 and later uses up to about 35 percent more tokens for the same text than earlier Claude models, varying by content.

The language. The same declaration took from 1.1× the English token count (Chinese (Simplified)) to 1.8× (Japanese) with o200k_base. Words per token also depends on how long a language’s words are: German and Turkish build long words from several parts, so they show few words per token even though the same text costs only 1.3× and 1.5× the English token count.

The kind of text. Code, Markdown and JSON are full of symbols, indentation and short fragments. With o200k_base, English prose came to 0.87 words per token, this site’s technical guides to 0.70, Markdown to 0.58, TypeScript to 0.47 and pretty-printed JSON to 0.25. Minified JSON has almost no spaces, so its “words” are meaningless: convert it by characters instead.

How we measured

For languages we used the Universal Declaration of Human Rights in 19 languages from the UDHR in Unicode project: the title, preamble and all 30 articles (1,747 words in English), without markup or notes. Because every version says the same thing, the token counts compare like with like. It is formal, legal language, so everyday text can come out a little different. For kinds of text we used English from this site: the reader-facing text of our 17 guides, 57 TypeScript source files, 10 Markdown documents, and our model price file as pretty-printed and minified JSON.

Words are counted the way word processors count them, by splitting on spaces. Chinese and Japanese are written without spaces between words, and splitting them into words needs a dictionary that varies between tools, so for those we convert by characters. Characters are Unicode characters without spaces, after normalising the text to the standard composed form (NFC) that almost all web text uses. Tokens were counted with OpenAI’s o200k_base and cl100k_base (through gpt-tokenizer, which matches OpenAI’s tiktoken) and with the official open-weight tokenizers for DeepSeek V4, Qwen 3.8 and Mistral Small 4, the same files our token counter uses. DeepSeek V4 Pro and V4.1 Flash gave identical counts on every sample, so they share a row. The measurement ran on 2026-10-11; the script is in the site’s source code, so anyone can repeat it.

Claude’s tokenizer isn’t public, so we can’t measure it. Instead the converter uses Anthropic’s published figure, that 1 million tokens is roughly 555,000 words on current models and 750,000 words on models before Claude Opus 4.7. It is labelled as Anthropic’s figure and only offered for English prose, because Anthropic gives no figures for other languages or for code. Gemini was counted with Google’s official countTokens API (gemini-3.8-flash).

Pages

A page isn’t a fixed unit. The converter uses a common convention for 12-point text with 1-inch margins: about 500 words single-spaced or 250 double-spaced. Printed books often hold fewer words per page, and slides far fewer, so treat page counts as rough. To see how many tokens fit a model, use the context window checker; to price them, use the LLM cost calculator.

Examples

Worked examples

English prose: a blog post, a 100-page document and two context windows

Tokenizer1,000-word blog post100-page document128K tokens holds1M tokens holds
OpenAI o200k_base1,155 tokens57,728 tokens110,866 words222 pages866,138 words1,732 pages
OpenAI cl100k_base1,154 tokens57,699 tokens110,921 words222 pages866,567 words1,733 pages
DeepSeek V41,145 tokens57,270 tokens111,752 words224 pages873,063 words1,746 pages
Qwen 3.81,197 tokens59,874 tokens106,891 words214 pages835,086 words1,670 pages
Mistral Small 41,178 tokens58,901 tokens108,657 words217 pages848,882 words1,698 pages
Gemini1,186 tokens59,302 tokens107,923 words216 pages843,147 words1,686 pages
Claude (Opus 4.7 and later)Anthropic’s figure1,802 tokens90,090 tokens71,040 words142 pages555,000 words1,110 pages
Claude (before Opus 4.7)Anthropic’s figure1,333 tokens66,667 tokens96,000 words192 pages750,000 words1,500 pages

Pages are single-spaced at 500 words; double-spaced pages (250 words) are twice as many. The 100-page document is the text only: if you upload a PDF file, some providers also count each page as an image, which adds tokens.

1,000 tokens in words, by language and tokenizer

Languageo200k_basecl100k_baseDeepSeek V4Qwen 3.8Mistral Small 4GeminiSame text vs English (o200k)
English8668678738358498431.0×
Spanish7806466817327307521.2×
French7406246516997266981.3×
German6444985456296286201.3×
Portuguese (Brazil)7525996597167007131.2×
Italian6325445926866496661.4×
Russian5703115245725215761.4×
Arabic5602544865725965101.2×
Hindi6311883464575407421.6×
Urdu6952504755576417291.5×
Bengali4231193453053545971.7×
Chinese (Simplified)characters, not words1,2448381,7421,7141,0871,4081.1×
Japanesecharacters, not words1,1508471,5151,7031,2581,6941.8×
Korean4322543904914844421.4×
Indonesian5564325266355135761.5×
Turkish4563423304394204611.5×
Vietnamese8114585328588288471.5×
Persian6282755545966476331.4×
Swahili5964374424484555171.4×

Words (or characters for Chinese and Japanese) per 1,000 tokens, measured on the Universal Declaration of Human Rights. The last column compares token counts for the same declaration, so it shows how much more the same meaning costs. Measured 2026-10-11.

1,000 tokens in words, by kind of English text

Texto200k_basecl100k_baseDeepSeek V4Qwen 3.8Mistral Small 4GeminiCharacters per token (o200k)
Prose8668678738358498434.41
Technical writing7026986866566626453.42
Source code4664654474254434033.00
JSON, pretty-printed2522502432102302102.18
JSON, minified1010109993.24
Markdown5825785705525535373.24

Prose is the English declaration; technical writing is this site’s 17 guides; code is 57 TypeScript files; Markdown is 10 documentation files; JSON is our 186,845-character model price file. Characters don’t count spaces.

FAQ

Frequently asked questions

How many words is 1,000 tokens?

About 866 words of English prose with OpenAI’s o200k_base tokenizer (GPT-5, GPT-4.1, GPT-4o), by our measurement on the Universal Declaration of Human Rights. The tokenizers we measured gave 835 to 873. Technical writing full of numbers and names came to 702, and Anthropic’s own figure for current Claude models is 555.

How many tokens is 1,000 words?

About 1,155 tokens of English prose with o200k_base, or 1,425 for technical writing with numbers and model names. For current Claude models, Anthropic’s figure of 0.555 words per token gives about 1,802. Code, JSON and most other languages take more tokens for the same number of words.

How many pages is 128K or 1 million tokens?

For English prose with o200k_base, 128,000 tokens is about 110,866 words, or 222 single-spaced pages of 500 words. One million tokens is about 866,138 words or 1,732 pages. With Claude’s current tokenizer, Anthropic puts 1 million tokens at about 555,000 words.

Why do newer Claude models use more tokens?

Every model family trains its own tokenizer. Anthropic says the tokenizer introduced with Claude Opus 4.7 uses roughly 1 to 1.35 times as many tokens for the same text as earlier Claude models, depending on the content: about 555,000 words per million tokens, against 750,000 before. Claude’s tokenizer isn’t public, so that figure is Anthropic’s, not our measurement.

Why do other languages use more tokens than English?

A tokenizer gives whole-word tokens to the text it saw most when it was built, usually English, so English words are often a single token while other languages are split into more pieces. The same declaration took 1.1× the English token count in Chinese (Simplified) and 1.8× in Japanese with o200k_base. Older vocabularies are worse: cl100k_base needed 3.3× as many tokens as o200k_base for Hindi.

How many characters is a token?

For English prose with o200k_base, about 5.27 characters including spaces, or 4.41 without, so OpenAI’s rule of thumb of about four characters per token is on the safe side for plain English. Technical writing came to 3.42 characters per token without spaces, TypeScript 3.00 and minified JSON 3.24. Chinese and Japanese are converted by characters, because they are written without spaces.

Why does code or JSON use so many tokens?

Code and JSON are dense with brackets, quotes, symbols and indentation, which tokenizers split into many short pieces, and their “words” are long runs of symbols. With o200k_base we measured 0.47 words per token for TypeScript and 0.25 for pretty-printed JSON, against 0.87 for prose. For JSON, characters are the better guide: about 3.24 characters per token minified.

How accurate is this converter?

The ratios are real measurements, but they are averages for one kind of text, and your text will differ: names, numbers, formatting and writing style all change the count. Use the converter for planning, and paste the actual text into the token counter when you need the exact number for a model.