How to count tokens
- Paste your prompt, system message or document. The count updates as you type.
- Read the token count. It is exact for OpenAI models that use the
o200k_baseencoding (the GPT-4o family and newer OpenAI models), computed by the same byte-pair encoding the API uses. - Choose a context window size to see whether the text fits and how much room is left for the model's answer.
- Open Estimate API cost, enter the per-million-token prices from your provider's pricing page and the expected output length, and get the cost per request and for a batch.
The first time you open this page, the tokenizer vocabulary (about 1 MB compressed) downloads once. Until it has loaded, a rough estimate is shown and labeled as such.
What is a token?
Language models don't read characters or words. They read tokens: chunks of text from a fixed vocabulary of around 200,000 entries for o200k_base. Common English words are usually one token, including the leading space (" hello"). Rare words, long numbers, code and many non-English scripts split into several tokens. The colored view above shows each token as a separate highlight, which makes it easy to see why some text is "expensive".
Examples (o200k_base)
Hello, world! | 4 tokens: Hello , world ! |
|---|---|
| English prose | Roughly 4 characters or ¾ of a word per token |
| JSON and code | Often fewer characters per token because of punctuation and indentation. Minifying JSON before sending it can save tokens. |
Accuracy and other model families
Counts are exact only for the o200k_base encoding. Older OpenAI models (GPT-4, GPT-3.5) use cl100k_base and give somewhat different counts. Anthropic Claude, Google Gemini, Meta Llama and Mistral each use their own tokenizers, so treat this number as a ballpark for them. The difference is usually modest for English prose but can be large for code and non-English text. For billing-grade numbers, use the provider's own token counting endpoint, such as the Anthropic token counting API or the Gemini countTokens method.
Real API requests also add a few tokens of overhead per message for chat formatting, and tools, images and function definitions count too. The cost estimate is therefore a planning figure, not an invoice.
Tips to reduce token usage
- Remove boilerplate instructions that repeat in every request, or use your provider's prompt caching for long static prefixes.
- Minify JSON and strip unneeded fields before including data in a prompt.
- Ask for concise output and set a maximum output length. Output tokens usually cost several times more than input tokens.
Frequently asked questions
Which models is this token count exact for?
OpenAI models that use the o200k_base encoding, which includes the GPT-4o family and newer OpenAI models. For other providers the number is an approximation.
Why aren't model prices built in?
Prices change often and differ by provider, region and tier. Hard-coded prices go stale quickly, so the estimator uses the numbers you enter from your provider's current pricing page.
Is my prompt sent to OpenAI or anywhere else?
No. Tokenization runs in your browser with an open-source implementation of the tokenizer (the gpt-tokenizer library, MIT license). No API key is used and no request is made with your text.
How many words is 1,000 tokens?
For typical English text, about 750 words. For code, other languages or text full of numbers and symbols it can be much fewer.
Why does the context window need room left over?
The context window covers both your input and the model's output. If your prompt uses 127,000 of 128,000 tokens, the answer can be at most 1,000 tokens long.
Related tools
Last updated: